<html xmlns:o="urn:schemas-microsoft-com:office:office" xmlns:w="urn:schemas-microsoft-com:office:word" xmlns:m="http://schemas.microsoft.com/office/2004/12/omml" xmlns="http://www.w3.org/TR/REC-html40"><head><meta http-equiv=Content-Type content="text/html; charset=utf-8"><meta name=Generator content="Microsoft Word 15 (filtered medium)"><style><!--
/* Font Definitions */
@font-face
        {font-family:"Cambria Math";
        panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
        {font-family:Calibri;
        panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
        {font-family:"Segoe UI Symbol";
        panose-1:2 11 5 2 4 2 4 2 2 3;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
        {margin:0in;
        font-size:10.0pt;
        font-family:"Calibri",sans-serif;}
a:link, span.MsoHyperlink
        {mso-style-priority:99;
        color:blue;
        text-decoration:underline;}
span.EmailStyle19
        {mso-style-type:personal-reply;
        font-family:"Calibri",sans-serif;
        color:windowtext;}
.MsoChpDefault
        {mso-style-type:export-only;
        font-size:10.0pt;
        mso-ligatures:none;}
@page WordSection1
        {size:8.5in 11.0in;
        margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
        {page:WordSection1;}
--></style></head><body lang=EN-US link=blue vlink=purple style='word-wrap:break-word'><div class=WordSection1><p class=MsoNormal><span style='font-size:11.0pt'>It seems to me that we are full circle back to a Turing Test.  If the LLM encodes and demonstrates skill (they certainly do), and these skills can be progress a solution of some real-world problem, then it is just empty chauvinism to say they don\u2019t understand a topic.<o:p></o:p></span></p><p class=MsoNormal><span style='font-size:11.0pt'><o:p>&nbsp;</o:p></span></p><div style='border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0in 0in 0in'><p class=MsoNormal style='margin-bottom:12.0pt'><b><span style='font-size:12.0pt;color:black'>From: </span></b><span style='font-size:12.0pt;color:black'>Friam &lt;friam-bounces@redfish.com&gt; on behalf of Steve Smith &lt;sasmyth@swcp.com&gt;<br><b>Date: </b>Thursday, September 11, 2025 at 10:12\u202fAM<br><b>To: </b>friam@redfish.com &lt;friam@redfish.com&gt;<br><b>Subject: </b>Re: [FRIAM] Hallucinations<o:p></o:p></span></p></div><div><p class=MsoNormal><span style='font-size:11.0pt'>I find LLM engagement to be somewhere between that with a highly <br>plausible gossip and a well researched survey paper in a subject I am <br>interested in?<br><br>Where a given conversation lands in this interval almost exclusively <br>seems to rely on my care in crafting my prompts.<br><br>I don't expect 'truth' out of either gossip or a survey paper... just <br>'perspective'?<br><br>On 9/11/25 10:55 am, glen wrote:<br>&gt; OK. You're right in principle. But we might want to think of this in <br>&gt; the context of all algorithms. For example, let's say you run a FFT on <br>&gt; a signal and it outputs some frequencies. Does the signal *actually* <br>&gt; contain or express those frequencies? Or is it just an inference that <br>&gt; we find reliable?<br>&gt;<br>&gt; The same is true of the LLM inferences. Whether one ascribes truth or <br>&gt; falsity to those inferences is only relevant to metaphysicians and <br>&gt; philosophers. What matters is how reliable the inferences are when we <br>&gt; do some task. Yelling at the kids on your lawn doesn't achieve <br>&gt; anything. It's better to go out there and talk to them. 8^D<br>&gt;<br>&gt;<br>&gt; On 9/10/25 8:38 PM, Russ Abbott wrote:<br>&gt;&gt; Glen, I wish people would stop talking about whether LLM-generated <br>&gt;&gt; sentences are true or false. The mechanisms LLMs employ to generate a <br>&gt;&gt; sentence have nothing to do with whether the sentence turns out to be <br>&gt;&gt; true or false. A sentence may have a higher probability of being true <br>&gt;&gt; if the training data consisted entirely of true sentences. (Even <br>&gt;&gt; that's not guaranteed; similar true sentences might have their <br>&gt;&gt; components interchanged when used during generation.) But the point <br>&gt;&gt; is: the transformer process has no connection to the validity&nbsp;of its <br>&gt;&gt; output. If an LLM reliably generates true sentences, no credit is due <br>&gt;&gt; to the transformer. If the training data consists entirely of <br>&gt;&gt; true/false sentences, the generated output is more likely to be <br>&gt;&gt; true/false. Output validity plays no role in how an LLM generates its <br>&gt;&gt; output.<br>&gt;&gt;<br>&gt;&gt; Marcus, if an LLM is trained entirely on false statements, its <br>&gt;&gt; &quot;confidence&quot; in its output will presumably be the same as it would be <br>&gt;&gt; if it were trained entirely on true statements. Truthfulness is not a <br>&gt;&gt; consideration in the generation process. Speaking of a need to reduce <br>&gt;&gt; ambiguity suggests that the LLM understands the input and realizes it <br>&gt;&gt; might have multiple meanings. But of course, LLMs don't understand <br>&gt;&gt; anything, they don't realize anything, and they can't take meaning <br>&gt;&gt; into consideration when generating output.<br>&gt;&gt;<br>&gt;&gt;<br>&gt;&gt;<br>&gt;&gt;<br>&gt;&gt;<br>&gt;&gt; On Tue, Sep 9, 2025 at 5:20\u202fPM glen &lt;gepropella@gmail.com <br>&gt;&gt; &lt;<a href="mailto:gepropella@gmail.com">mailto:gepropella@gmail.com</a>&gt;&gt; wrote:<br>&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp; It's unfortunate jargon [</span><span style='font-size:11.0pt;font-family:"Segoe UI Symbol",sans-serif'>\u26e7</span><span style='font-size:11.0pt'>]. So it's nothing like whether an LLM <br>&gt;&gt; is red (unless you adopt a jargonal definition of &quot;red&quot;). And your <br>&gt;&gt; example is a great one for understanding how language fluency *is* at <br>&gt;&gt; least somewhat correlated with fidelity. The statistical probability <br>&gt;&gt; of the phrase &quot;LLMs hallucinate&quot; is &gt;&gt; 0, whereas the prob for the <br>&gt;&gt; phrase &quot;LLMs are red&quot; is vanishingly small. It would be the same for <br>&gt;&gt; black swans and Lewis Carroll writings *if* they weren't canonical <br>&gt;&gt; teaching devices. It can't be that sophisticated if children think <br>&gt;&gt; it's funny.<br>&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp; But imagine all the woo out there where words like &quot;entropy&quot; or <br>&gt;&gt; &quot;entanglement&quot; are used falsely. IDK for sure, but my guess is the <br>&gt;&gt; false sentences outnumber the true ones by a lot. So the LLM has a <br>&gt;&gt; high probability of forming false sentences.<br>&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp; Of course, in that sense, if a physicist finds themselves talking <br>&gt;&gt; to an expert in the &quot;Law of Attraction&quot; (e.g. the movie &quot;The Secret&quot;) <br>&gt;&gt; and makes scientifically true statements about entanglement, the guru <br>&gt;&gt; may well judge them as false. So there's &quot;true in context&quot; (validity) <br>&gt;&gt; and &quot;ontologically true&quot; (soundness). A sentence can be true in <br>&gt;&gt; context but false in the world and vice versa, depending on who's in <br>&gt;&gt; control of the reinforcement.<br>&gt;&gt;<br>&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp; [</span><span style='font-size:11.0pt;font-family:"Segoe UI Symbol",sans-serif'>\u26e7</span><span style='font-size:11.0pt'>] We could discuss the strength of the analogy between human <br>&gt;&gt; hallucination and LLM &quot;hallucination&quot;, especially in the context of <br>&gt;&gt; prediction coding. But we don't need to. Just consider it jargon and <br>&gt;&gt; move on.<br>&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp; On 9/9/25 4:37 PM, Russ Abbott wrote:<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; Marcus, Glen,<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; Your responses are much too sophisticated for me. Now that I'm <br>&gt;&gt; retired (and, in truth, probably before as well), I tend to think in <br>&gt;&gt; much simpler terms.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; My basic point was to express my surprise at realizing that it <br>&gt;&gt; makes as much sense to ask whether an LLM hallucinates as it does to <br>&gt;&gt; ask whether an LLM is red. It's a category&nbsp;mismatch--at least I now <br>&gt;&gt; think so.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; _<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; _<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; __-- Russ &lt;https://russabbott.substack.com/ <br>&gt;&gt; &lt;<a href="https://russabbott.substack.com/">https://russabbott.substack.com/</a>&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt; On Tue, Sep 9, 2025 at 3:45\u202fPM glen &lt;gepropella@gmail.com <br>&gt;&gt; &lt;<a href="mailto:gepropella@gmail.com">mailto:gepropella@gmail.com</a>&gt; &lt;mailto:gepropella@gmail.com <br>&gt;&gt; &lt;<a href="mailto:gepropella@gmail.com">mailto:gepropella@gmail.com</a>&gt;&gt;&gt; wrote:<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp;The question of whether fluency is (well) correlated to <br>&gt;&gt; accuracy seems to assume something like mentalizing, the idea that <br>&gt;&gt; there's a correspondence between minds mediated by a correspondence <br>&gt;&gt; between the structure of the world and the structure of our <br>&gt;&gt; minds/language. We've talked about the &quot;interface theory of <br>&gt;&gt; perception&quot;, where Hoffman (I think?) argues we're more likely to <br>&gt;&gt; learn *false* things than we are true things. And we've argued about <br>&gt;&gt; realism, pragmatism, prediction coding, and everything else under the <br>&gt;&gt; sun on this list.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp;So it doesn't surprise me if most people assume there will <br>&gt;&gt; be more true statements in the corpus than false statements, at least <br>&gt;&gt; in domains where there exists a common sense, where the laity *can* <br>&gt;&gt; perceive the truth. In things like quantum mechanics or whatever, <br>&gt;&gt; then all bets are off becuase there are probably more false sentences <br>&gt;&gt; than true ones.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp;If there are more true than false sentences in the corpus, <br>&gt;&gt; then reinforcement methods like Marcus' only bear a small burden (in <br>&gt;&gt; lay domains). The implicit fidelity does the lion's share. But in <br>&gt;&gt; those domains where counter-intuitive facts dominate, the <br>&gt;&gt; reinforcement does the most work.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp;On 9/9/25 3:12 PM, Marcus Daniels wrote:<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; Three ways some to mind..&nbsp; I would guess that OpenAI, <br>&gt;&gt; Google, Anthropic, and xAI are far more sophisticated..<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;&nbsp; 1. Add a softmax penalty to the loss that tracks <br>&gt;&gt; non-factual statements or grammatical constraints. &nbsp;Cross entropy may <br>&gt;&gt; not understand that some parts of content are more important than <br>&gt;&gt; others.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;&nbsp; 2. Change how the beam search works during inference <br>&gt;&gt; to skip sequences that fail certain predicates \u2013 like a lookahead <br>&gt;&gt; that says \u201cOh, I can\u2019t say that..\u201d<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;&nbsp; 3. Grade the output, either using human or non-LLM <br>&gt;&gt; supervision, and re-train.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; *From:*Friam &lt;friam-bounces@redfish.com <br>&gt;&gt; &lt;<a href="mailto:friam-bounces@redfish.com">mailto:friam-bounces@redfish.com</a>&gt; &lt;mailto:friam-bounces@redfish.com <br>&gt;&gt; &lt;<a href="mailto:friam-bounces@redfish.com">mailto:friam-bounces@redfish.com</a>&gt;&gt;&gt; *On Behalf Of *Russ Abbott<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; *Sent:* Tuesday, September 9, 2025 3:03 PM<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; *To:* The Friday Morning Applied Complexity Coffee <br>&gt;&gt; Group &lt;friam@redfish.com &lt;<a href="mailto:friam@redfish.com">mailto:friam@redfish.com</a>&gt; <br>&gt;&gt; &lt;mailto:friam@redfish.com &lt;<a href="mailto:friam@redfish.com">mailto:friam@redfish.com</a>&gt;&gt;&gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; *Subject:* [FRIAM] Hallucinations<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; OpenAI just published a paper on hallucinations <br>&gt;&gt; &lt;https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf <br>&gt;&gt; &lt;<a href="https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf">https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf</a>&gt; <br>&gt;&gt; &lt;https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf <br>&gt;&gt; &lt;<a href="https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf">https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf</a>&gt;&gt;&gt;&nbsp;as <br>&gt;&gt; well as a post summarizing the paper <br>&gt;&gt; &lt;https://openai.com/index/why-language-models-hallucinate/ <br>&gt;&gt; &lt;<a href="https://openai.com/index/why-language-models-hallucinate/">https://openai.com/index/why-language-models-hallucinate/</a>&gt; <br>&gt;&gt; &lt;https://openai.com/index/why-language-models-hallucinate/ <br>&gt;&gt; &lt;<a href="https://openai.com/index/why-language-models-hallucinate/">https://openai.com/index/why-language-models-hallucinate/</a>&gt;&gt;&gt;. The <br>&gt;&gt; two of them seem wrong-headed in such a simple and obvious way that <br>&gt;&gt; I'm surprised the issue they discuss is still alive.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; The paper and post point out that LLMs are trained to <br>&gt;&gt; generate fluent language--which they do extraordinarily well. The <br>&gt;&gt; paper and post also point out that LLMs are not trained to <br>&gt;&gt; distinguish valid from invalid statements. Given those facts about <br>&gt;&gt; LLMs, it's not clear why one should expect LLMs to be able to <br>&gt;&gt; distinguish true statements from false statements--and hence why one <br>&gt;&gt; should expect to be able to prevent LLMs from hallucinating.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; In other words, LLMs are built to generate text; they <br>&gt;&gt; are not built to understand the texts they generate and certainly not <br>&gt;&gt; to be able to determine whether the texts they generate make <br>&gt;&gt; factually correct or incorrect statements.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; Please see my post <br>&gt;&gt; &lt;https://russabbott.substack.com/p/why-language-models-hallucinate-according <br>&gt;&gt; &lt;<a href="https://russabbott.substack.com/p/why-language-models-hallucinate-according">https://russabbott.substack.com/p/why-language-models-hallucinate-according</a>&gt; <br>&gt;&gt; &lt;https://russabbott.substack.com/p/why-language-models-hallucinate-according <br>&gt;&gt; &lt;<a href="https://russabbott.substack.com/p/why-language-models-hallucinate-according">https://russabbott.substack.com/p/why-language-models-hallucinate-according</a>&gt;&gt;&gt; <br>&gt;&gt; elaborating on this.<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt; Why is this not obvious, and why is OpenAI still <br>&gt;&gt; talking about it?<br>&gt;&gt; &nbsp;&nbsp;&nbsp;&nbsp; &gt;&nbsp; &nbsp; &nbsp; &gt;<br>&gt;&gt; &nbsp;&nbsp;&nbsp; -- <br>&gt;<br>&gt;<br><br>.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ... --- -- . / .- .-. . / ..- ... . ..-. ..- .-..<br>FRIAM Applied Complexity Group listserv<br>Fridays 9a-12p Friday St. Johns Cafe&nbsp;&nbsp; /&nbsp;&nbsp; Thursdays 9a-12p Zoom <a href="https://bit.ly/virtualfriam">https://bit.ly/virtualfriam</a><br>to (un)subscribe <a href="http://redfish.com/mailman/listinfo/friam_redfish.com">http://redfish.com/mailman/listinfo/friam_redfish.com</a><br>FRIAM-COMIC <a href="http://friam-comic.blogspot.com/">http://friam-comic.blogspot.com/</a><br>archives:&nbsp; 5/2017 thru present <a href="https://redfish.com/pipermail/friam_redfish.com/">https://redfish.com/pipermail/friam_redfish.com/</a><br>&nbsp; 1/2003 thru 6/2021&nbsp; <a href="http://friam.383.s1.nabble.com/">http://friam.383.s1.nabble.com/</a><o:p></o:p></span></p></div></div></body></html>