I’ve been contemplating this situation. I think that what the author wants out of a Jev-like model is not at all what I want out of it.
> I decided to check this on questions where the answer is well understood. For example:
> A classical particle of mass m is embedded in a system at thermodynamic equilibrium with temperature T. What is its velocity v?
If I feed that into a model, the answer I want is: “the combination of the model and the provided state has nothing useful to add to your prior”.
If I want to know the Maxwell-Boltzmann distribution, I can look it up or I can derive it or I can ask a fancy LLM to do it for me (at the cost of some reasoning tokens and some time - unless I’m using an ultraspeed inference system, I’m not getting this answer in 50ms). [0]
Similarly, if I want to know that 73% of incoming customer support requests are spam/fraud, I should measure that - it’s a property of my system, it takes some manual classification and a database query, and it will be a different percentage than your customer support system would see. I neither expect nor want my classifier to know this (unless I’m using a conventional classifier manually trained on my data, and the whole point of Jev is to avoid this).
What I want out of a system like Jev is to tell me how the probabilities change as a result of the per-sample data I provide. Which, is the case of this Boltzmann distribution question, is nothing: I provided no data and the classifier can infer nothing.
[0] A really good answer would observe that the answer depends on the dimension of the system (probably 3, but 2D systems are a thing) and also on whether the particles are hot enough for relativistic effects to matter (probably not). And maybe a good answer would check whether the material is a gas - the answer for a solid is not the same, but I suppose that’s not classical. Oh, and one shouldn’t forget drift: if you have a classical particle in a moving fluid or a classical charged particle in an electric field, you will again get a different answer.
Yes, I’m being pedantic. But if you want good answers you should be pedantic, and the Jev-like model is not where the pedantry should go.
The point of the theoretical problems is that they should be the easiest cases to handle. How can you trust the probabilities from real world classifiers if it can't even handle well defined problems.
1. It’s ridiculous. Not only is the question prima facie absurd for a model of this type, it sort of doesn’t fit into the whole training model. An LLM (charitably) predicts token probabilities, which one might generalize to mean that the LLM operates on probability distributions over strings. So asking for the probability of “fraud” versus “not fraud” makes sense. But asking for the probability of “1.23” versus “2.7” is kind of out of distribution - those are numbers, no one is training on an entire continuum of two-decimal-place real numbers, and similar numbers can have wildly different representations (“25.4” vs “25.40” vs “2.54e1”).
2. It’s barely a classification problem as written. If I wanted it to be a classification problem, maybe I would try:
“There is a machine that receives little sealed containers of air. In each container one molecule is painted red. The machine measured the interior temperature of one particular container and determined that it was 300K.” Question: in what range was the velocity of the red molecule at the instant that the container entered the machine. Choices: 0-100m/s, 100-200m/s, etc.
I maintain that this question is a weird thing to train a Jev-like model on and that I don’t think that a classifier I use would need to answer it well.
I do find it disappointing that Jev conflates “the probabilities are all equal” with “I have no clue”, and I think it would be better if it were at least clearly documented how the model’s ability to figure something out relates to the API response probability.
I think this comment misunderstands the nature of a probabilistic system. It doesn't reason, or use constraints, or do analytic math. It measures probabilities based on observations.
The well-defined problems aren't well-defined in this sense.
> So Jev does actually know which distributions are correct, it just fails to produce them
Claims like this needs to be deeply analyized. Models do have some emergent capabilities[1], and I think there's a lot of evidence to show that semantics is actually learned (Word2Vec), and some math seems like it also might be learned (e.g. modular arithmetic). But it's hard to exactly say where there's some internal mechanism generating a true answer and where we're just getting lucky with some distribution so the answer just seems right.
IMO by asking Jev underspecified questions like this, you're essentially using it as a random number generator (similar to the dice example). On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated: https://leonardgrazian.com/blog/jev-calibration/
In the dice example[0] and in this one, the expected output is a distribution e.g. "1: 16.7%, 2: 16.7%...", not a random generation. I agree that in implementation, the architecture is not designed to accurately calculate or incorporate any known uncertainties like these, but the task itself is completely reasonable and arguably trivial.
IMO the lesson here is: even in trivial cases, Jev's outputs are just ~reasonableness scores which do not correspond to actual probabilities. They should not be treated as actual probabilities without careful calibration and plenty of meta-uncertainty about how well that calibration extrapolates.
The problem is, most of the value proposition of Jev is that it gives you the probabilities without doing that, which it doesn't.
Jev should ideally respond in a non-random way, it should just list out the probabilities.
I suppose trying to interface with the model like this is like asking an LLM how many times the letter E appears in a word - just not the correct way to ask that given its model
Wouldn't asking humans have this same kind of problem too?
If you asked a set of humans to generate a random distribution of heads or tails from coin flips, it wouldn't be similar to a real world distribution of coin flips either because we also have our own biases. [1]
That's interesting - unless the numbers are all made up, this didn't read like slop to me. Somebody is claiming to have done this and that and concluded something. Maybe it's your slopatron that's miscalibrated! Have you considered jev :-D
I’ve been contemplating this situation. I think that what the author wants out of a Jev-like model is not at all what I want out of it.
> I decided to check this on questions where the answer is well understood. For example:
> A classical particle of mass m is embedded in a system at thermodynamic equilibrium with temperature T. What is its velocity v?
If I feed that into a model, the answer I want is: “the combination of the model and the provided state has nothing useful to add to your prior”.
If I want to know the Maxwell-Boltzmann distribution, I can look it up or I can derive it or I can ask a fancy LLM to do it for me (at the cost of some reasoning tokens and some time - unless I’m using an ultraspeed inference system, I’m not getting this answer in 50ms). [0]
Similarly, if I want to know that 73% of incoming customer support requests are spam/fraud, I should measure that - it’s a property of my system, it takes some manual classification and a database query, and it will be a different percentage than your customer support system would see. I neither expect nor want my classifier to know this (unless I’m using a conventional classifier manually trained on my data, and the whole point of Jev is to avoid this).
What I want out of a system like Jev is to tell me how the probabilities change as a result of the per-sample data I provide. Which, is the case of this Boltzmann distribution question, is nothing: I provided no data and the classifier can infer nothing.
[0] A really good answer would observe that the answer depends on the dimension of the system (probably 3, but 2D systems are a thing) and also on whether the particles are hot enough for relativistic effects to matter (probably not). And maybe a good answer would check whether the material is a gas - the answer for a solid is not the same, but I suppose that’s not classical. Oh, and one shouldn’t forget drift: if you have a classical particle in a moving fluid or a classical charged particle in an electric field, you will again get a different answer.
Yes, I’m being pedantic. But if you want good answers you should be pedantic, and the Jev-like model is not where the pedantry should go.
The point of the theoretical problems is that they should be the easiest cases to handle. How can you trust the probabilities from real world classifiers if it can't even handle well defined problems.
Two reasons:
1. It’s ridiculous. Not only is the question prima facie absurd for a model of this type, it sort of doesn’t fit into the whole training model. An LLM (charitably) predicts token probabilities, which one might generalize to mean that the LLM operates on probability distributions over strings. So asking for the probability of “fraud” versus “not fraud” makes sense. But asking for the probability of “1.23” versus “2.7” is kind of out of distribution - those are numbers, no one is training on an entire continuum of two-decimal-place real numbers, and similar numbers can have wildly different representations (“25.4” vs “25.40” vs “2.54e1”).
2. It’s barely a classification problem as written. If I wanted it to be a classification problem, maybe I would try:
“There is a machine that receives little sealed containers of air. In each container one molecule is painted red. The machine measured the interior temperature of one particular container and determined that it was 300K.” Question: in what range was the velocity of the red molecule at the instant that the container entered the machine. Choices: 0-100m/s, 100-200m/s, etc.
I maintain that this question is a weird thing to train a Jev-like model on and that I don’t think that a classifier I use would need to answer it well.
I do find it disappointing that Jev conflates “the probabilities are all equal” with “I have no clue”, and I think it would be better if it were at least clearly documented how the model’s ability to figure something out relates to the API response probability.
I think this comment misunderstands the nature of a probabilistic system. It doesn't reason, or use constraints, or do analytic math. It measures probabilities based on observations.
The well-defined problems aren't well-defined in this sense.
> So Jev does actually know which distributions are correct, it just fails to produce them
Claims like this needs to be deeply analyized. Models do have some emergent capabilities[1], and I think there's a lot of evidence to show that semantics is actually learned (Word2Vec), and some math seems like it also might be learned (e.g. modular arithmetic). But it's hard to exactly say where there's some internal mechanism generating a true answer and where we're just getting lucky with some distribution so the answer just seems right.
[1] https://arxiv.org/pdf/2502.00873
IMO by asking Jev underspecified questions like this, you're essentially using it as a random number generator (similar to the dice example). On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated: https://leonardgrazian.com/blog/jev-calibration/
In the dice example[0] and in this one, the expected output is a distribution e.g. "1: 16.7%, 2: 16.7%...", not a random generation. I agree that in implementation, the architecture is not designed to accurately calculate or incorporate any known uncertainties like these, but the task itself is completely reasonable and arguably trivial.
IMO the lesson here is: even in trivial cases, Jev's outputs are just ~reasonableness scores which do not correspond to actual probabilities. They should not be treated as actual probabilities without careful calibration and plenty of meta-uncertainty about how well that calibration extrapolates.
The problem is, most of the value proposition of Jev is that it gives you the probabilities without doing that, which it doesn't.
[0] https://kantahayashiai.github.io/posts/jev-does-not-play-dic...
Jev should ideally respond in a non-random way, it should just list out the probabilities.
I suppose trying to interface with the model like this is like asking an LLM how many times the letter E appears in a word - just not the correct way to ask that given its model
Wouldn't asking humans have this same kind of problem too?
If you asked a set of humans to generate a random distribution of heads or tails from coin flips, it wouldn't be similar to a real world distribution of coin flips either because we also have our own biases. [1]
[1] "Heads or tails?"--a reachability bias in binary choice" - https://pubmed.ncbi.nlm.nih.gov/24773285/
I don't think so. The author is expecting Jev to output the equivalent of "H: 50%, T:50%", not the equivalent of "TTHTTHTHH..."
Isn't this just testing its training on statistical mechanics? It's sort of a specialized field of knowledge and perhaps not well trained on that.
It might be better at classifying the author from a piece of text, given N candidates. That's a dirac delta function, unless it's been plagiarized.
Some interesting data from RH about these zero proof guardrails: https://developers.redhat.com/articles/2026/10/02/benchmarki...
this is related to, but not the same as, the question i had in the original thread: https://news.ycombinator.com/item?id=49719097
Also, it's a very small and dumb model. That's the only reason why it's so fast.
not very
[flagged]
Can you please not post AI-generated or AI-edited comments to HN? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079.
Of course, it's impossible to know for sure what was LLM processed or not, but this post got classified that way.
That's interesting - unless the numbers are all made up, this didn't read like slop to me. Somebody is claiming to have done this and that and concluded something. Maybe it's your slopatron that's miscalibrated! Have you considered jev :-D