<html xmlns:v="urn:schemas-microsoft-com:vml" xmlns:o="urn:schemas-microsoft-com:office:office" xmlns:w="urn:schemas-microsoft-com:office:word" xmlns:m="http://schemas.microsoft.com/office/2004/12/omml" xmlns="http://www.w3.org/TR/REC-html40"><head><meta http-equiv=Content-Type content="text/html; charset=utf-8"><meta name=Generator content="Microsoft Word 15 (filtered medium)"><!--[if !mso]><style>v\:* {behavior:url(#default#VML);}
o\:* {behavior:url(#default#VML);}
w\:* {behavior:url(#default#VML);}
.shape {behavior:url(#default#VML);}
</style><![endif]--><style><!--
/* Font Definitions */
@font-face
{font-family:"Cambria Math";
panose-1:2 4 5 3 5 4 6 3 2 4;}
@font-face
{font-family:Calibri;
panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
{font-family:Tahoma;
panose-1:2 11 6 4 3 5 4 4 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
{margin:0in;
font-size:10.0pt;
font-family:"Calibri",sans-serif;}
a:link, span.MsoHyperlink
{mso-style-priority:99;
color:blue;
text-decoration:underline;}
span.m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18
{mso-style-name:m_-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18;}
span.EmailStyle19
{mso-style-type:personal-reply;
font-family:"Calibri",sans-serif;
color:windowtext;}
.MsoChpDefault
{mso-style-type:export-only;
font-size:10.0pt;
mso-ligatures:none;}
@page WordSection1
{size:8.5in 11.0in;
margin:1.0in 1.0in 1.0in 1.0in;}
div.WordSection1
{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext="edit" spidmax="1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext="edit">
<o:idmap v:ext="edit" data="1" />
</o:shapelayout></xml><![endif]--></head><body lang=EN-US link=blue vlink=purple style='word-wrap:break-word'><div class=WordSection1><p class=MsoNormal><span style='font-size:11.0pt'>There is a significant literature on these topics\u2026<br><br><a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">https://transformer-circuits.pub/2024/scaling-monosemanticity/</a><br><a href="https://arxiv.org/abs/1702.01135">https://arxiv.org/abs/1702.01135</a><br><a href="https://arxiv.org/html/2502.02470v2#abstract">https://arxiv.org/html/2502.02470v2#abstract</a><o:p></o:p></span></p><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p><div style='border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0in 0in 0in'><p class=MsoNormal style='margin-bottom:12.0pt'><b><span style='font-size:12.0pt;color:black'>From: </span></b><span style='font-size:12.0pt;color:black'>Russ Abbott <russ.abbott@gmail.com><br><b>Date: </b>Saturday, September 13, 2025 at 5:46\u202fPM<br><b>To: </b>Marcus Daniels <marcus@snoutfarm.com><br><b>Cc: </b>The Friday Morning Applied Complexity Coffee Group <friam@redfish.com><br><b>Subject: </b>Re: [FRIAM] Hallucinations<o:p></o:p></span></p></div><div><div><p class=MsoNormal><span style='font-size:12.0pt;font-family:"Arial",sans-serif;color:black'>I wasn't expecting the LLM process to parallel the human process. I want to know what LLMs produce as an embedding and what that embedding might allow them to do. This is all about what an LLM can do on its own (without external tools). This isn't a new question. I'm surprised there hasn't been more work along these lines.<o:p></o:p></span></p></div><div><p class=MsoNormal><span style='font-size:11.0pt'><br clear=all><o:p></o:p></span></p></div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p></div><div><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p></div><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p></div><p class=MsoNormal><span style='font-size:11.0pt'><o:p> </o:p></span></p><div><div><p class=MsoNormal><span style='font-size:11.0pt'>On Fri, Sep 12, 2025 at 7:31\u202fPM Marcus Daniels <<a href="mailto:marcus@snoutfarm.com">marcus@snoutfarm.com</a>> wrote:<o:p></o:p></span></p></div><blockquote style='border:none;border-left:solid #CCCCCC 1.0pt;padding:0in 0in 0in 6.0pt;margin-left:4.8pt;margin-right:0in'><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;margin-bottom:12.0pt'><span style='font-size:11.0pt'>Understanding how language and logic become encoded in deep neural nets is interesting topic. However, we don\u2019t have that expectation of one another. \u201cYour argument is plausible, but have your synaptic connectivity (a bunch of floats) been imaged and studied for correctness?\u201d <o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>If one wants to be sure that reasoning is sound, relying on human or LLM intuition is the wrong way to go about it. Instead, with MCP one can leverage LLM language skills to formulate (and refine & run) code that can show that the proofs that they generate (or hallucinate) are sound. <o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div style='border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0in 0in 0in'><p class=MsoNormal style='mso-margin-top-alt:auto;margin-bottom:12.0pt'><b><span style='font-size:12.0pt;color:black'>From: </span></b><span style='font-size:12.0pt;color:black'>Russ Abbott <<a href="mailto:russ.abbott@gmail.com">russ.abbott@gmail.com</a>><br><b>Date: </b>Friday, September 12, 2025 at 7:16\u202fPM<br><b>To: </b>Marcus Daniels <<a href="mailto:marcus@snoutfarm.com">marcus@snoutfarm.com</a>><br><b>Cc: </b>The Friday Morning Applied Complexity Coffee Group <<a href="mailto:friam@redfish.com">friam@redfish.com</a>><br><b>Subject: </b>Re: [FRIAM] Hallucinations</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>I'm not sure what your point is. You would expect Lean to be able to do that. Also, my example didn't say that the combination of statements is invalid. You added that explicitly and asked Lean to confirm it.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>My interest is what the encoding of the natural-language input looks like and--since it just looks like a sequence of floats--what it conveys in the context of that particular learned embedding framework.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>-- Russ<o:p></o:p></span></p></div></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>On Fri, Sep 12, 2025, 5:04\u202fPM Marcus Daniels <<a href="mailto:marcus@snoutfarm.com">marcus@snoutfarm.com</a>> wrote:<o:p></o:p></span></p></div><blockquote style='border:none;border-left:solid #CCCCCC 1.0pt;padding:0in 0in 0in 6.0pt;margin-left:4.8pt;margin-top:5.0pt;margin-right:0in;margin-bottom:5.0pt'><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>Let\u2019s have Claude formulate the contradiction in Lean 4 and delegate the reasoning to a tool that is good at that. <br>(Just like I wouldn\u2019t do long division by hand.)<br><br><img border=0 width=604 height=583 style='width:6.2916in;height:6.0729in' id="m_-1232332303123083915m_-7859194655259392957Picture_x0020_1" src="cid:image001.png@01DC2407.560DD3B0"><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div style='border:none;border-top:solid #E1E1E1 1.0pt;padding:3.0pt 0in 0in 0in'><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><b><span style='font-size:11.0pt'>From:</span></b><span style='font-size:11.0pt'> Russ Abbott <<a href="mailto:russ.abbott@gmail.com">russ.abbott@gmail.com</a>> <br><b>Sent:</b> Friday, September 12, 2025 3:51 PM<br><b>To:</b> Marcus Daniels <<a href="mailto:marcus@snoutfarm.com">marcus@snoutfarm.com</a>><br><b>Cc:</b> The Friday Morning Applied Complexity Coffee Group <<a href="mailto:friam@redfish.com">friam@redfish.com</a>><br><b>Subject:</b> Re: [FRIAM] Hallucinations<o:p></o:p></span></p></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>Marcus, You're right, and I was wrong. I was much too insistent that LLMs don't understand the text they manipulate.</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>A couple of weeks ago, I asked ChatGPT to embed (encode) a sentence and then decode it back to natural language. It said it didn't have access to the tools to do exactly that, but it would show me what the result would look like.</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>The input sentence was: \u201cI ate an apple because I was hungry. The apple was rotten. I got sick. My friend ate a banana. The banana was not rotten. My friend didn\u2019t get sick.\u201d</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>ChatGPT simulated embedding/encoding the sentence as a vector. It then produced what it claimed was a reasonable natural language approximation of that vector. The result was: "A person and their friend ate fruit. One of the fruits was rotten, which caused sickness, while the other was fresh and did not cause illness."</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>If ChatGPT can be believed, this is quite impressive. It implies that the embedding/encoding of natural language text includes something like the essential semantics of the original text. I had forgotten all about this when I wrote my post about hallucinations. I apologize.</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>What I would like to do now -- and perhaps someone can help figure out if any tools are available to do this -- is to explore more carefully the sorts of information embeddings/encodings contain. For example, what would one get if one encoded and then decoded Chomsky's famous sentence: "Colorless green ideas sleep furiously." What would one get if one encoded -> decoded a contradiction: "All men are mortal. Socrates is a man; Socrates is immortal." What about: "The integer 3 is larger than the integer 9." Or "The American Revolutionary War occurred during the 19th century. George Washington led the American troops in that war. </span><span style='font-size:11.0pt'>George Washington's tenure as the inaugural president of the United States began on April 30, 1789." Etc.<o:p></o:p></span></p></div></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>-- <a href="https://russabbott.substack.com/">Russ Abbott</a> (Click for my Substack)<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>Professor Emeritus, Computer Science<br>California State University, Los Angeles<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>On Thu, Sep 11, 2025 at 6:59</span><span style='font-size:11.0pt;font-family:"Arial",sans-serif'>\u202f</span><span style='font-size:11.0pt'>PM Marcus Daniels <<a href="mailto:marcus@snoutfarm.com">marcus@snoutfarm.com</a>> wrote:<o:p></o:p></span></p></div><blockquote style='border:none;border-left:solid #CCCCCC 1.0pt;padding:0in 0in 0in 6.0pt;margin-left:4.8pt;margin-top:5.0pt;margin-right:0in;margin-bottom:5.0pt'><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt;font-family:"Tahoma",sans-serif'>\ufeff</span></span><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'>It often works with the frontier models to take a computational science or theory paper and to have them implement the idea expressed in some computer language. One can also often invert that program back into natural language (and/or with LaTeX equations). Further, one can translate between very different formal languages (imperative vs. functional), which would be hard work for most people.</span></span><span style='font-size:11.0pt'><br><br><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18>These summaries and transformations work so well that tools like Github Copilot will periodically perform a conversation summary, and simply drop the whole conversation and start over with crystallized context (due to context window limitations). When it picks up after that, one will often see a few syntax or API misunderstandings before it regroups to where it was.</span><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'> </span></span><span style='font-size:11.0pt'><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'>What this pivoting ease implies to me is that LLMs have a deep semantic representation of the conversation (and knowledge and these skills). It certainly is not just a matter of mating token sequences with some deft smoothing.</span></span><span style='font-size:11.0pt'><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'> </span></span><span style='font-size:11.0pt'><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'>Another example that has come up for me recently is using LLMs to predict simulation or solver outputs. When faced with learning large arrays of numbers, what it does is more like capturing a picture then a sequence of digits. It doesn\u2019t know, without some help, about why number boundaries, signs, and decimal points are important. Only through hard-won experience does it learn that the most and least significant digits should be treated differently. Syntax is a hint one can offer through weak scaffolding penalties (outside of the training material). It learns the semantics first. Strong syntax penalties can get in the way of learning semantics by creating problematic energy barriers.</span></span><span style='font-size:11.0pt'><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'> </span></span><span style='font-size:11.0pt'><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span class=m-1232332303123083915m-7859194655259392957m-6590922116472562574emailstyle18><span style='font-size:11.0pt'>While LLMs are huge, the Chinchilla optimality criterion (20 tokens per parameter) forces regularization. There\u2019s some flood fill, but I don\u2019t think it can hold up for idiosyncratic lexical patterns.</span></span><span style='font-size:11.0pt'><o:p></o:p></span></p><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div style='border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0in 0in 0in'><p class=MsoNormal style='mso-margin-top-alt:auto;margin-bottom:12.0pt'><b><span style='font-size:11.0pt;color:black'>From: </span></b><span style='font-size:11.0pt;color:black'>Friam <<a href="mailto:friam-bounces@redfish.com">friam-bounces@redfish.com</a>> on behalf of Santafe <<a href="mailto:desmith@santafe.edu">desmith@santafe.edu</a>><br><b>Date: </b>Thursday, September 11, 2025 at 5:12</span><span style='font-size:11.0pt;font-family:"Arial",sans-serif;color:black'>\u202f</span><span style='font-size:11.0pt;color:black'>PM<br><b>To: </b><a href="mailto:Russ.Abbott@gmail.com">Russ.Abbott@gmail.com</a> <<a href="mailto:Russ.Abbott@gmail.com">Russ.Abbott@gmail.com</a>>, The Friday Morning Applied Complexity Coffee Group <<a href="mailto:friam@redfish.com">friam@redfish.com</a>><br><b>Subject: </b>Re: [FRIAM] Hallucinations</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>In your post, Russ, you say: <o:p></o:p></span></p><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>\u201c</span><span style='font-size:14.5pt;color:#363737;background:white'>They are trained to produce fluent language, not to produce valid statements.</span><span style='font-size:11.0pt'>\u201c <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>Is that actually, operationally, what they are trained to do? I speak from a position of ignorance here, but my impression is that they are trained to effectively stitch together fragments of varying lengths, according to rules for what stitchings are \u201ccompatible\u201d.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>My thinking here is metaphorical, to homologous recombination in DNA. Some regions that don\u2019t start out contiguous can be concatenated by DNA repair machinery, because under the physics to which it responds, they have plausible enough overlap that it considers them \u201ccompatible\u201d or eligible to be identified at the join region, their \u201cmis-matches\u201d edited out. Other pairs are so dissimilar that, under its operating physics, the repair machinery will effectively never join them.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>My metaphor isn\u2019t great, in the sense that if what LLMs (for human speech) are doing is \u201cnext-word prediction\u201d, that says that the notion of \u201cjoining\u201d is reduced formally to appending next-words onto strings. Though, to the extent that certain substrings of next-words are extremely frequently attested across the corpus of all the training expressions, one would expect to see extended sequences essentially reproduced as fragments with large probability.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>If my previous two characterizations aren\u2019t fundamentally wrong, it would follow that fluent speech-generation becomes possible because the compatible-joining relations are suffficiently strong in human languages that the attention structures or other feed-forward aspects of the architecture have no trouble capturing them in parameters, even though human linguists trying to write them as re-write rules from which a computer could generate native-like speech failed for decades to get anywhere close to that. My interpretation here would be consistent with what I believed was the main watershed change in the LLMs: that the parametric models would, ultimately, have terribly few parameters, whereas the LLMs can flood-fill a corpus with parameters, and then try to drip out the parts that don\u2019t \u201cstick to\u201d some pattern in the data, and are regarded as the excess entropy from the sampling algorithm that the training is supposed to recognize and remove. It is easy to imagine that fluent speech has far more regularities than rule-book linguists captured parametrically, but still few enough that LLMs can have no trouble attaching to almost-all of them, with parameters to spare. Hence fluent speech could be epiphenomenal on what they are (operationally, mechanistically) being trained to do, but a natural summary statistic for the effectiveness of that training, and of course the one that drives market engagement.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>But if the above is the case, then the question of when they get \u201cthe syntax\u201d right and \u201cthe semantics\u201d wrong, would seem to turn on how much context from the training set is needed to identify semantically as well as syntactically appropriate \u201callowed joins\u201d of fragments. When short fragments contain enough of their own context to constrain most of the semantics, the stitching training algorithm has no reason to perform any worse at revealing the semantic signal in the training set than the syntactic one. But if probability needs to be withheld for a long time in the prediction model, driving it to prioritize a much smaller number of longer or more remote assembled inputs from the training data, it could still do fine on syntax but fail to \u201cfind\u201d and \u201crender\u201d the semantic signal in the training data, even if that signal is present in principal. <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>I would not feel a need to use terms like \u201cunderstanding\u201d anywhere in the above, to make predictions of what kinds of successes or failures an LLM might deliver from the user\u2019s perspective. It seems to me like something that all lives in the domain of hardness-of-search combinatorics in data-spaces with a lot of difficult structure.<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>Eric<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;margin-bottom:12.0pt'><span style='font-size:11.0pt'> <o:p></o:p></span></p><blockquote style='margin-top:5.0pt;margin-bottom:5.0pt'><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>On Sep 10, 2025, at 7:02, Russ Abbott <<a href="mailto:russ.abbott@gmail.com">russ.abbott@gmail.com</a>> wrote:<o:p></o:p></span></p></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><div><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'>OpenAI just published a <a href="https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fcdn.openai.com%2fpdf%2fd04913be-3f6f-4d2b-b283-ff432ef4aaa5%2fwhy-language-models-hallucinate.pdf&c=E,1,IvBfvLzhn3L6LCNk3_ktKoEbc9NI2Oqq8vlFpNcIXCHElptIB-Fx-UxQYyTnCFW_ToeD5Kd4RjHkY-6fLxSBqZueOcvRqyHwpsHPK9ugMNcsOw,,&typo=1">paper on hallucinations</a> as well as <a href="https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fopenai.com%2findex%2fwhy-language-models-hallucinate%2f&c=E,1,tEcctM28Lbt5XBi3gNiUX-RiFelMYHNq6K3VJBilGv1_Z8uAt34ta8FaU-FcW5i8V3-2tsjNPu_at8Es78G2_drdmykgOltvjRvvaw1hUgnXUsv3&typo=1">a post summarizing the paper</a>. The two of them seem wrong-headed in such a simple and obvious way that I'm surprised the issue they discuss is still alive. </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'>The paper and post point out that LLMs are trained to generate fluent language--which they do extraordinarily well. The paper and post also point out that LLMs are not trained to distinguish valid from invalid statements. Given those facts about LLMs, it's not clear why one should expect LLMs to be able to distinguish true statements from false statements--and hence why one should expect to be able to prevent LLMs from hallucinating. </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'>In other words, LLMs are built to generate text; they are not built to understand the texts they generate and certainly not to be able to determine whether the texts they generate make factually correct or incorrect statements.</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'>Please see <a href="https://linkprotect.cudasvc.com/url?a=https%3a%2f%2frussabbott.substack.com%2fp%2fwhy-language-models-hallucinate-according&c=E,1,zLy4H6KEpD5hDchYiBUjiH2J5dG2O9bmqa-jm1z6mGgRSqZgDKaVd2D2Xh_2Wuzi7FtZu2kjIOTNjQuk4iwsnfNUG68UPCxmZvD_IHTVUEPTcW6HDgpmcozzRQ,,&typo=1">my post</a> elaborating on this.</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'> </span><span style='font-size:11.0pt'><o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt;font-family:"Arial",sans-serif'>Why is this not obvious, and why is OpenAI still talking about it?</span><span style='font-size:11.0pt'><o:p></o:p></span></p></div></div></div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>-- <a href="https://linkprotect.cudasvc.com/url?a=https%3a%2f%2frussabbott.substack.com%2f&c=E,1,cKMmq0etz4RiUaE4G2F04re6Su0EnNyqR9j5Dx8RcccQVNOB2r5CMNBzxRL9EYmN3lG_11nhB4wP-5jPf7NR86Mb9VxP9Jn2YUdKPZQT&typo=1">Russ Abbott</a> (Click for my Substack)<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>Professor Emeritus, Computer Science<br>California State University, Los Angeles<o:p></o:p></span></p></div><div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'>.- .-.. .-.. / ..-. --- --- - . .-. ... / .- .-. . / .-- .-. --- -. --. / ... --- -- . / .- .-. . / ..- ... . ..-. ..- .-..<br>FRIAM Applied Complexity Group listserv<br>Fridays 9a-12p Friday St. Johns Cafe / Thursdays 9a-12p Zoom <a href="https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fbit.ly%2fvirtualfriam&c=E,1,TGWUFxByQV3GAAU3oSRoMNfDJD6ptWzY73PWkEy6wjvRSnx8Mc4UYZvwNnCZtQTtnx4s1YQWhA5OFZgcHYsPOfh2UOY3Y08aOLzFbRROXd4isiXdoT93L5Ncgw,,&typo=1">https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fbit.ly%2fvirtualfriam&c=E,1,TGWUFxByQV3GAAU3oSRoMNfDJD6ptWzY73PWkEy6wjvRSnx8Mc4UYZvwNnCZtQTtnx4s1YQWhA5OFZgcHYsPOfh2UOY3Y08aOLzFbRROXd4isiXdoT93L5Ncgw,,&typo=1</a><br>to (un)subscribe <a href="https://linkprotect.cudasvc.com/url?a=http%3a%2f%2fredfish.com%2fmailman%2flistinfo%2ffriam_redfish.com&c=E,1,R8rvP64Y8Ojn7C4RmXsVaTwfI61-h--86QYAcdZfJB5b2Vma9UVdbCXCsDqLzWtC_TM9Ckm5LlRcoIn4_6mGC8c_WptkWvx_WtZA0PdtE8ViiUc,&typo=1">https://linkprotect.cudasvc.com/url?a=http%3a%2f%2fredfish.com%2fmailman%2flistinfo%2ffriam_redfish.com&c=E,1,R8rvP64Y8Ojn7C4RmXsVaTwfI61-h--86QYAcdZfJB5b2Vma9UVdbCXCsDqLzWtC_TM9Ckm5LlRcoIn4_6mGC8c_WptkWvx_WtZA0PdtE8ViiUc,&typo=1</a><br>FRIAM-COMIC <a href="https://linkprotect.cudasvc.com/url?a=http%3a%2f%2ffriam-comic.blogspot.com%2f&c=E,1,JolQcZ2iD8sPfKhQE-npSBJtUmqqa8EaE0J19wBCnesx4rjYKUpByO5mwjwVUiEn91veQr1Bk3B0gvLNuTtgIkN8-2VZRSkQS61pFh_zro8Oe_g7&typo=1">https://linkprotect.cudasvc.com/url?a=http%3a%2f%2ffriam-comic.blogspot.com%2f&c=E,1,JolQcZ2iD8sPfKhQE-npSBJtUmqqa8EaE0J19wBCnesx4rjYKUpByO5mwjwVUiEn91veQr1Bk3B0gvLNuTtgIkN8-2VZRSkQS61pFh_zro8Oe_g7&typo=1</a><br>archives: 5/2017 thru present <a href="https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fredfish.com%2fpipermail%2ffriam_redfish.com%2f&c=E,1,3J72bQm1T2SIdCaPyxSx4gitJ3Bt_OjLNAoKxcLa4u2f5Yw2m3gHImwAjCKE9RabMTMbzMedGiltpwWQ5w10fnNmDFvVkW9oQcfwHVezCQ,,&typo=1">https://linkprotect.cudasvc.com/url?a=https%3a%2f%2fredfish.com%2fpipermail%2ffriam_redfish.com%2f&c=E,1,3J72bQm1T2SIdCaPyxSx4gitJ3Bt_OjLNAoKxcLa4u2f5Yw2m3gHImwAjCKE9RabMTMbzMedGiltpwWQ5w10fnNmDFvVkW9oQcfwHVezCQ,,&typo=1</a><br> 1/2003 thru 6/2021 <a href="http://friam.383.s1.nabble.com/">http://friam.383.s1.nabble.com/</a><o:p></o:p></span></p></div></blockquote></div><p class=MsoNormal style='mso-margin-top-alt:auto;mso-margin-bottom-alt:auto'><span style='font-size:11.0pt'> <o:p></o:p></span></p></div></div></div></div></blockquote></div></div></div></blockquote></div></div></div></div></blockquote></div></div></body></html>