<html xmlns:o="urn:schemas-microsoft-com:office:office" xmlns:w="urn:schemas-microsoft-com:office:word" xmlns:m="http://schemas.microsoft.com/office/2004/12/omml" xmlns="http://www.w3.org/TR/REC-html40"><head><meta http-equiv=Content-Type content="text/html; charset=utf-8"><meta name=Generator content="Microsoft Word 15 (filtered medium)"><style><!--
/* Font Definitions */
@font-face
{font-family:"Cambria Math";
panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
{font-family:Calibri;
panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
{font-family:"Segoe UI Symbol";
panose-1:2 11 5 2 4 2 4 2 2 3;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
{margin:0in;
font-size:10.0pt;
font-family:"Calibri",sans-serif;}
a:link, span.MsoHyperlink
{mso-style-priority:99;
color:blue;
text-decoration:underline;}
span.EmailStyle19
{mso-style-type:personal-reply;
font-family:"Calibri",sans-serif;
color:windowtext;}
.MsoChpDefault
{mso-style-type:export-only;
font-size:10.0pt;
mso-ligatures:none;}
@page WordSection1
{size:8.5in 11.0in;
margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
{page:WordSection1;}
--></style></head><body lang=EN-US link=blue vlink=purple style='word-wrap:break-word'><div class=WordSection1><p class=MsoNormal><span style='font-size:11.0pt'>It seems to me that we are full circle back to a Turing Test. If the LLM encodes and demonstrates skill (they certainly do), and these skills can be progress a solution of some real-world problem, then it is just empty chauvinism to say they don\u2019t understand a topic.<o:p></o:p></span></p><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p><div style='border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0in 0in 0in'><p class=MsoNormal style='margin-bottom:12.0pt'><b><span style='font-size:12.0pt;color:black'>From: </span></b><span style='font-size:12.0pt;color:black'>Friam <friam-bounces@redfish.com> on behalf of Steve Smith <sasmyth@swcp.com><br><b>Date: </b>Thursday, September 11, 2025 at 10:12\u202fAM<br><b>To: </b>friam@redfish.com <friam@redfish.com><br><b>Subject: </b>Re: [FRIAM] Hallucinations<o:p></o:p></span></p></div><div><p class=MsoNormal><span style='font-size:11.0pt'>I find LLM engagement to be somewhere between that with a highly <br>plausible gossip and a well researched survey paper in a subject I am <br>interested in?<br><br>Where a given conversation lands in this interval almost exclusively <br>seems to rely on my care in crafting my prompts.<br><br>I don't expect 'truth' out of either gossip or a survey paper... just <br>'perspective'?<br><br>On 9/11/25 10:55 am, glen wrote:<br>> OK. You're right in principle. But we might want to think of this in <br>> the context of all algorithms. For example, let's say you run a FFT on <br>> a signal and it outputs some frequencies. Does the signal *actually* <br>> contain or express those frequencies? Or is it just an inference that <br>> we find reliable?<br>><br>> The same is true of the LLM inferences. Whether one ascribes truth or <br>> falsity to those inferences is only relevant to metaphysicians and <br>> philosophers. What matters is how reliable the inferences are when we <br>> do some task. Yelling at the kids on your lawn doesn't achieve <br>> anything. It's better to go out there and talk to them. 8^D<br>><br>><br>> On 9/10/25 8:38 PM, Russ Abbott wrote:<br>>> Glen, I wish people would stop talking about whether LLM-generated <br>>> sentences are true or false. The mechanisms LLMs employ to generate a <br>>> sentence have nothing to do with whether the sentence turns out to be <br>>> true or false. A sentence may have a higher probability of being true <br>>> if the training data consisted entirely of true sentences. (Even <br>>> that's not guaranteed; similar true sentences might have their <br>>> components interchanged when used during generation.) But the point <br>>> is: the transformer process has no connection to the validity of its <br>>> output. If an LLM reliably generates true sentences, no credit is due <br>>> to the transformer. If the training data consists entirely of <br>>> true/false sentences, the generated output is more likely to be <br>>> true/false. Output validity plays no role in how an LLM generates its <br>>> output.<br>>><br>>> Marcus, if an LLM is trained entirely on false statements, its <br>>> "confidence" in its output will presumably be the same as it would be <br>>> if it were trained entirely on true statements. Truthfulness is not a <br>>> consideration in the generation process. Speaking of a need to reduce <br>>> ambiguity suggests that the LLM understands the input and realizes it <br>>> might have multiple meanings. But of course, LLMs don't understand <br>>> anything, they don't realize anything, and they can't take meaning <br>>> into consideration when generating output.<br>>><br>>><br>>><br>>><br>>><br>>> On Tue, Sep 9, 2025 at 5:20\u202fPM glen <gepropella@gmail.com <br>>> <<a href="mailto:gepropella@gmail.com">mailto:gepropella@gmail.com</a>>> wrote:<br>>><br>>> It's unfortunate jargon [</span><span style='font-size:11.0pt;font-family:"Segoe UI Symbol",sans-serif'>\u26e7</span><span style='font-size:11.0pt'>]. So it's nothing like whether an LLM <br>>> is red (unless you adopt a jargonal definition of "red"). And your <br>>> example is a great one for understanding how language fluency *is* at <br>>> least somewhat correlated with fidelity. The statistical probability <br>>> of the phrase "LLMs hallucinate" is >> 0, whereas the prob for the <br>>> phrase "LLMs are red" is vanishingly small. It would be the same for <br>>> black swans and Lewis Carroll writings *if* they weren't canonical <br>>> teaching devices. It can't be that sophisticated if children think <br>>> it's funny.<br>>><br>>> But imagine all the woo out there where words like "entropy" or <br>>> "entanglement" are used falsely. IDK for sure, but my guess is the <br>>> false sentences outnumber the true ones by a lot. So the LLM has a <br>>> high probability of forming false sentences.<br>>><br>>> Of course, in that sense, if a physicist finds themselves talking <br>>> to an expert in the "Law of Attraction" (e.g. the movie "The Secret") <br>>> and makes scientifically true statements about entanglement, the guru <br>>> may well judge them as false. So there's "true in context" (validity) <br>>> and "ontologically true" (soundness). A sentence can be true in <br>>> context but false in the world and vice versa, depending on who's in <br>>> control of the reinforcement.<br>>><br>>><br>>> [</span><span style='font-size:11.0pt;font-family:"Segoe UI Symbol",sans-serif'>\u26e7</span><span style='font-size:11.0pt'>] We could discuss the strength of the analogy between human <br>>> hallucination and LLM "hallucination", especially in the context of <br>>> prediction coding. But we don't need to. Just consider it jargon and <br>>> move on.<br>>><br>>> On 9/9/25 4:37 PM, Russ Abbott wrote:<br>>> > Marcus, Glen,<br>>> ><br>>> > Your responses are much too sophisticated for me. Now that I'm <br>>> retired (and, in truth, probably before as well), I tend to think in <br>>> much simpler terms.<br>>> ><br>>> > My basic point was to express my surprise at realizing that it <br>>> makes as much sense to ask whether an LLM hallucinates as it does to <br>>> ask whether an LLM is red. It's a category mismatch--at least I now <br>>> think so.<br>>> > _<br>>> > _<br>>> > __-- Russ <https://russabbott.substack.com/ <br>>> <<a href="https://russabbott.substack.com/">https://russabbott.substack.com/</a>>><br>>> ><br>>> ><br>>> ><br>>> ><br>>> > On Tue, Sep 9, 2025 at 3:45\u202fPM glen <gepropella@gmail.com <br>>> <<a href="mailto:gepropella@gmail.com">mailto:gepropella@gmail.com</a>> <mailto:gepropella@gmail.com <br>>> <<a href="mailto:gepropella@gmail.com">mailto:gepropella@gmail.com</a>>>> wrote:<br>>> ><br>>> > The question of whether fluency is (well) correlated to <br>>> accuracy seems to assume something like mentalizing, the idea that <br>>> there's a correspondence between minds mediated by a correspondence <br>>> between the structure of the world and the structure of our <br>>> minds/language. We've talked about the "interface theory of <br>>> perception", where Hoffman (I think?) argues we're more likely to <br>>> learn *false* things than we are true things. And we've argued about <br>>> realism, pragmatism, prediction coding, and everything else under the <br>>> sun on this list.<br>>> ><br>>> > So it doesn't surprise me if most people assume there will <br>>> be more true statements in the corpus than false statements, at least <br>>> in domains where there exists a common sense, where the laity *can* <br>>> perceive the truth. In things like quantum mechanics or whatever, <br>>> then all bets are off becuase there are probably more false sentences <br>>> than true ones.<br>>> ><br>>> > If there are more true than false sentences in the corpus, <br>>> then reinforcement methods like Marcus' only bear a small burden (in <br>>> lay domains). The implicit fidelity does the lion's share. But in <br>>> those domains where counter-intuitive facts dominate, the <br>>> reinforcement does the most work.<br>>> ><br>>> ><br>>> > On 9/9/25 3:12 PM, Marcus Daniels wrote:<br>>> > > Three ways some to mind.. I would guess that OpenAI, <br>>> Google, Anthropic, and xAI are far more sophisticated..<br>>> > ><br>>> > > 1. Add a softmax penalty to the loss that tracks <br>>> non-factual statements or grammatical constraints. Cross entropy may <br>>> not understand that some parts of content are more important than <br>>> others.<br>>> > > 2. Change how the beam search works during inference <br>>> to skip sequences that fail certain predicates \u2013 like a lookahead <br>>> that says \u201cOh, I can\u2019t say that..\u201d<br>>> > > 3. Grade the output, either using human or non-LLM <br>>> supervision, and re-train.<br>>> > ><br>>> > > *From:*Friam <friam-bounces@redfish.com <br>>> <<a href="mailto:friam-bounces@redfish.com">mailto:friam-bounces@redfish.com</a>> <mailto:friam-bounces@redfish.com <br>>> <<a href="mailto:friam-bounces@redfish.com">mailto:friam-bounces@redfish.com</a>>>> *On Behalf Of *Russ Abbott<br>>> > > *Sent:* Tuesday, September 9, 2025 3:03 PM<br>>> > > *To:* The Friday Morning Applied Complexity Coffee <br>>> Group <friam@redfish.com <<a href="mailto:friam@redfish.com">mailto:friam@redfish.com</a>> <br>>> <mailto:friam@redfish.com <<a href="mailto:friam@redfish.com">mailto:friam@redfish.com</a>>>><br>>> > > *Subject:* [FRIAM] Hallucinations<br>>> > ><br>>> > > OpenAI just published a paper on hallucinations <br>>> <https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf <br>>> <<a href="https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf">https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf</a>> <br>>> <https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf <br>>> <<a href="https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf">https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf</a>>>> as <br>>> well as a post summarizing the paper <br>>> <https://openai.com/index/why-language-models-hallucinate/ <br>>> <<a href="https://openai.com/index/why-language-models-hallucinate/">https://openai.com/index/why-language-models-hallucinate/</a>> <br>>> <https://openai.com/index/why-language-models-hallucinate/ <br>>> <<a href="https://openai.com/index/why-language-models-hallucinate/">https://openai.com/index/why-language-models-hallucinate/</a>>>>. The <br>>> two of them seem wrong-headed in such a simple and obvious way that <br>>> I'm surprised the issue they discuss is still alive.<br>>> > ><br>>> > > The paper and post point out that LLMs are trained to <br>>> generate fluent language--which they do extraordinarily well. The <br>>> paper and post also point out that LLMs are not trained to <br>>> distinguish valid from invalid statements. Given those facts about <br>>> LLMs, it's not clear why one should expect LLMs to be able to <br>>> distinguish true statements from false statements--and hence why one <br>>> should expect to be able to prevent LLMs from hallucinating.<br>>> > ><br>>> > > In other words, LLMs are built to generate text; they <br>>> are not built to understand the texts they generate and certainly not <br>>> to be able to determine whether the texts they generate make <br>>> factually correct or incorrect statements.<br>>> > ><br>>> > > Please see my post <br>>> <https://russabbott.substack.com/p/why-language-models-hallucinate-according <br>>> <<a href="https://russabbott.substack.com/p/why-language-models-hallucinate-according">https://russabbott.substack.com/p/why-language-models-hallucinate-according</a>> <br>>> <https://russabbott.substack.com/p/why-language-models-hallucinate-according <br>>> <<a href="https://russabbott.substack.com/p/why-language-models-hallucinate-according">https://russabbott.substack.com/p/why-language-models-hallucinate-according</a>>>> <br>>> elaborating on this.<br>>> > ><br>>> > > Why is this not obvious, and why is OpenAI still <br>>> talking about it?<br>>> > ><br>>> -- <br>><br>><br><br>.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ... --- -- . / .- .-. . / ..- ... . ..-. ..- .-..<br>FRIAM Applied Complexity Group listserv<br>Fridays 9a-12p Friday St. Johns Cafe / Thursdays 9a-12p Zoom <a href="https://bit.ly/virtualfriam">https://bit.ly/virtualfriam</a><br>to (un)subscribe <a href="http://redfish.com/mailman/listinfo/friam_redfish.com">http://redfish.com/mailman/listinfo/friam_redfish.com</a><br>FRIAM-COMIC <a href="http://friam-comic.blogspot.com/">http://friam-comic.blogspot.com/</a><br>archives: 5/2017 thru present <a href="https://redfish.com/pipermail/friam_redfish.com/">https://redfish.com/pipermail/friam_redfish.com/</a><br> 1/2003 thru 6/2021 <a href="http://friam.383.s1.nabble.com/">http://friam.383.s1.nabble.com/</a><o:p></o:p></span></p></div></div></body></html>