Hacker Newsnew | past | comments | ask | show | jobs | submit | prometheus1992's commentslogin

Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).

I would assume this is a proxy for general intelligence. A model that can drive a car and do a bunch of other real world stuff is closer to a generalized intelligence that can reason through any task.

Tesla isn't using a general purpose model, they're using many highly-specialized models for a more deterministic system than "hey chat drive this car for me"

Why not? Benchmark all the things!

couldn't be more wrong - there are so many zero shot classifiers available on HF which do the same thing.

I think you're possibly arguing a point I wasn't making? I'm not saying Jev invented zero-shot classification, or that there aren't already zero-shot classifiers on HF that can do classification without fine-tuning; I was responding to questions asked in a silo.

>With BERT, you need a large, labeled dataset, and you have to train/fine-tune the model.

I think they were responding to this. You can use BERT to provide zero shot classification predictions.


I guess I assumed they meant BERT, not some specific BERT-base model. Vanilla BERT does not support zero shot.

it really only takes about 5 minutes of exposure to get the vitamin D your body needs. i would happily do it between 7-9AM when the UV is low than doing it at mid day.

You only need 5 minutes at midday in the summer. At 7am, you get essentially no vitamin D.

What you are saying is exactly the wrong takeaway. Please read the article:

https://onlinelibrary.wiley.com/doi/10.1111/j.1751-1097.2008...

Beneficial UV is UVB, which is low when the sun is low because it can’t get through the ozone and atmosphere as well when the sun is at an angle. Instead, you are getting mostly UVA, which has an easier time passing through the skin (and atmosphere) and doesn’t stimulate vitamin D production.


>>The state currently generates 62% of its electricity from renewable and zero-carbon sources, with the largest percentage of that coming from solar

62% is pretty respectable for the size of California (same electricity consumption as Spain). I did not know this.


And as of the writing of this comment, ~7pm, about 25% of CA's grid is supplied by batteries: https://www.caiso.com/todays-outlook/supply

I was shocked when I saw how much batteries contribute, honestly don't hear about them much but they have scaled a lot in recent years.


Amazing what California can achieve when they don't have to deal with constant cease-and-desist lawsuits, eminent domain battles, environmental review challenges, etc (crying in high speed rail)

…and massive subsidies for the developers at the federal level…

https://www.energy.ca.gov/data-reports/energy-almanac/califo...

Flattening the Duck Curve: batteries reach 44% of evening demand in California - https://news.ycombinator.com/item?id=47633698 - April 2026 (5 comments)


You might be surprised to know that worldwide, low carbon electricity production is a little above 40%

https://ourworldindata.org/grapher/electricity-mix?tab=line&...


Also the 2nd highest electricity prices in the USA behind Hawaii

Imagine how much more expensive it would be if renewables didn't make up such a big share.

Probably cheaper? Renewables have a big upfront Capex usually financed.

Everything has a big upfront Capex. However you finance it and because renewables are cheap fuel they come out cheap in the long run.

A large part of the Capex of renewables is land costs - which are high, but once you have the land you keep it. This cost is generally all financed as part of the initial build, so in 30 years everything is paid for - but you have to finance to replace only the renewables while keeping the land and so the second round is a lot cheaper.


For fossil fuels the cost in in the fuel, usually.

No, renewables are much cheaper.

Including the cost of the batteries to make it stable over the day? Would need a source for that. My understanding is it comparable at best, but certainly not "much cheaper".

Yes, I was impressed with this number as well. I wonder when we'd reach 75-80%, but I believe the zero-carbon electricity target is 2045?

[flagged]


grid maintenance is massive but wildfire liability is a larger problem California faces more than any other state

if only California didn't ban cultural burns in the first place


> and zero-carbon sources

this part of the claim is carrying a ton of weight here... Zero-carbon sources could include carbon dependent sources with paper offsets.


Electricity sources are tracked by the California Energy Commission, the 2024 report [0] shows the zero-carbon sources that make up this claim are:

  Nuclear     9.92%
  Large Hydro 11.06%
  Biomass     1.94%
  Geothermal  4.60%
  Small Hydro 1.52%
  Solar       21.30%
  Wind        11.89%
2% from biomass is suspect as a "zero-carbon" source, but the rest seems to be reasonably be zero carbon by today's standards.

[0] https://www.energy.ca.gov/data-reports/energy-almanac/califo...


Biomass is zero carbon. It pulled all the carbon it will emit from the air already. Burning it simply returns what it had already recently removed.

One can even potentially pyrolize the biomass and get biochar, giving up some energy but making it carbon negative power source.


> Biomass is zero carbon.

Not correct. Biomass includes burning wood from new growth forests. If the trees were kept in the ground they would be carbon sinks rather than carbon stores, therefore the overall effect is an increase in emissions. Furthermore, there's energy expended in prepping the wood (e.g. turning the trees into wood pellets ready to burn) and in transporting the wood to the biomass power station.


[flagged]


As someone who isn't in California, it's understandable that you might think this. It is, fortunately, wrong.

The biggest driver of costs of electricity in CA are (1) paying for damages caused by grid-initiated wildfires, and (2) paying to upgrade the grid to prevent future cases of (1).

This information comes directly from the government who approves rate increases for regulated utilities, and who has to publish their spending. PG&E, the primary power company, is also a public company and their financials are therefore public.

Rooftop solar specifically (and not general grid-scale renewables) do shift the cost balance of fixed-cost infra and consumption based usage, but the wildfire issues dwarf this.


For time of use rates ~0.34/kWh is the cheapest rate for overnight hours. 4-9pm is over 50c/kWh. It's crazy how prices have increased the last decade.

At least at those rates getting solar on your roof and a battery in the basement sounds like a bargain.

Why would OAI need to follow Jev? I really think this is paid by Jev. Jev itself won't have lunch money in a shortwhile because there are literally 10s of free alternatives available which can be run locally on basic consumer hardware. Terrible utility aside, there is no sensible business proposition in Jev.

I'm having a really hard time wrapping my head around why Jev is getting so much hype. It feels manufactured to me. I don't think they've proven a significant market for their product, and there's no independent benchmarks that prove anything. To me that's doesn't pass the smell test.

But if they had to show how well their product worked they might give away the whole game... because they'd have to compare their "noul" class against an NLI benchmark for instance, and possibly show they're losing to cross encoders and give away the fact that they are just rebranding NLI. Or rerankers (choice) or zero-shot classifiers.


The AI hype cycle is always looking for the next big thing. It really doesn't take much for enthusiasts to get very excited and push something into the stratosphere. Just not having a vibe coded website, and someone that made ChatGPT is enough to set them apart. Hitting pain points like pricing and speed and also implicitly mentioning llms (even if it's to mention it can't generate text in contrast to llms) make it seem like a major step-up.


They really don't match the performance of Jev. You can probably fine tune them to work well for a particular use case (assuming you have sufficient data).

But Jev is works pretty much out of the box without any fine tuning.


How are you forcing Qwen to answer in a structured way? I like this one better - https://github.com/deepanwadhwa/OpenDecision

I can tell you did not use AN llm to write this post. Thanks.

try this model on HF - https://huggingface.co/MoritzLaurer/deberta-v3-large-zerosho...

a 6 sided die rolled a 3

possible class names - the number is odd, the number is even

result:

the number is odd 0.945 the number is even 0.055

as someone else said, that 0.055 is probably bc of 6 and 3 being there.


@prodigycorp - reading your comments here on this post - you seem pretty hurt by this post.

Yeah, the reason why I am annoyed by it is because a person (who felt like a burner account of the laya creator) yesterday was haranguing me for saying that projects like this were vibe coded, posting the link to this project.

I evaluated this project yesterday and found its claims un-credible. It's literally nothing like jev. That's some context behind why, a day later, I find it annoying that this is somehow the top story on HN.

https://news.ycombinator.com/item?id=49752902


I don't really see the breakthrough in Jev. Classification, scoring, routing and returning probabilities over predefined choices are all established problems. We implemented category routing in our own retrieval system in a slightly different way: embed the incoming query, compare it against category profiles and route to the highest cosine-similarity. Obviously Jev isn't similarity based, but the underlying task of making a constrained decision from predefined choices isn't novel. TypeSafe says Jev has a new architecture and RLCD training, but Jev's actual architecture, weights and training details aren't public. So we can't even claim Jev is specifically a BERT classifier, but also don't see enough public technical evidence yet to call the underlying idea a breakthrough. Atleast they should publish a technical paper to prove their idea is breakthrough.

I agree with your analysis based on my own last night (on another HN post to this same gripe on reddit, before this blog post). OP received a lot of echo chamber support in the subreddit, and recommended to post to HN, so here we are.

The work is very amateurish, the "paper" would be a strong reject if I were still peer reviewing.

https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i_liter...


What is jev like? Did they release any research paper? I really think typesafe hired someone to boost their post because there was nothing "Breakthrough" about their product. At least this post has some touch with the reality that this functionality was available a year ago and was well known among ML people.

> At least this post has some touch with the reality that this functionality was available a year ago and was well known among ML people.

When you market a product you make exciting claims relative to the audience you’re engaging with. When was the last time you saw a product marketing page reverently lost all the academic research and prior art that came together to make a product possible?

If Layla’s functionality was available in a SaaS form in a way that could be used by all the people who are excited about and using Jev, wouldn’t this research have won hearts and minds last year when it landed? I would have a lot more empathy for the author if they’d taken a product to market and nobody cared. But even then maybe the market wasn’t ready. There are still reasonable explanations why sometimes ideas take off. We’re on a venture capital forum this shouldn’t need an explanation.


I can't believe you say in another post that you have experience with bert and yet you don't understand the value of a generalist classifier.

Good models take time and effort. There wasn't a good option for satisficers until a few days ago.


It is standard discourse on here if you look backwards. Attention is All You Need sounds like a big nothingburger according to this post: https://news.ycombinator.com/item?id=15938082

Did you see the carnage that typesafe's landing page was? every other post here is llm generated, every other poster here seems like a LLM.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: