Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

As I have been saying for years:

Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer. You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.

 help



For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today. Sometimes such articles interpret or place context around their hidden references, but a lot of the time they just summarize.

Giving someone the text output of a LLM is very similar to publishing a summary without links to the referenced material. When you were querying your LLM, you could have asked specific questions or asked for a custom focus or point of view. Your intended audience might have questions or different concerns, but they're unable to interact with your LLM. What you have delivered is static and unresponsive. It has all the disadvantages of being machine output without the advantage of being interactive, the way your LLM was for you.

It may have to wait until compute is cheap enough that tokens are essentially free, but we need a system to pass "hyperlinks" to LLM's primed with context, ready to be interactively queried on a chosen context. It's being overly generous to assume that people are putting even 300 bits into a LLM for every 1000 bits of regurgitated writing they try to pass off as their own. When people post LLM output as if it were their own, I have no choice but to assume they had zero knowledge of the subject, but this query taught them what they wanted to learn, and now they're sharing that. That's fine, but please pass an interactive LLM link rather than static text.

Once we have "hyperlinks" for LLM sessions, perhaps we can share LLM output a little more usefully and honestly.


I've seen professional journal pieces refer to science journal articles only to go and read the original article and find that it draws a different conclusion than what is implied by the journalist.

My father used to complain about that 40 years ago, though in his case he was reading newspapers rather than professional journals. But he'd point it out to me often enough that I started to see the pattern. Scientist publishes paper saying "We may have found evidence of X, which suggests the possibility that Y may also be occurring". Journalist: "Scientists find X which proves Y".

This has been happening for decades; I still see it happening today*. My cynical suspicion is that words like "maybe" and "suggests the possibility" don't sell enough papers.

* Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote". Summary said "Exposure to X can, on average, cause a 40% higher chance of Y" (where Y was a negative health outcome). I clicked through to the study and read it. Turned out the confidence interval on that chance of Y was so wide, all you could say with 95% confidence was that exposure to X could do anything from reduce your chance of Y by 5 percent, or increase it by 85 percent, or somewhere in between. They had averaged -5 and +85 to get the scarier-sounding 40% number that they published in the summary, but the truth would have been far closer to "this confidence interval is so wide that we really can't conclude anything from this data". But that wouldn't be nearly as likely to get them grants, so they tortured the data in their summary so that it would look better.


Scientist: "My discoveries are useless when taken out of context"

Media: "Scientists claim their discoveries are useless"


In most cases, they are only useful to other scientists in their field, which does mean that they are useless to the general person until they show up in the form of a new commodity.

There is also incentives to adapt the message to the outlet. If you send that data to peer review and say we found 40% increase, the reviewers would reject it, so they have to moderate themselves. But if you send a summary to the university’s outreach outlet saying that we found something or nothing we don’t know, then they would also reject putting it out. So even for the authors, the incentive is to send a careful conclusion to peer review and an overblown one to popsci.

Btw, I also think a 95% confidence interval is just the wrong statistic to look at given that data, and that they could probably have analyzed it better.


> Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote".

It's a similar thing. We live in the "attention economy", and research institutions - particularly after the US President openly went and had his minions cut funding to research purely on ideological reasons, but it's been a problem for decades - are just as susceptible to blow stuff out of proportion to make headlines and thus increase the chance someone might throw some money over the fence.

And media does the same, just to manufacture artificial debate. And so do politicians.

And frankly, I'm fed up with that, we will drive ourselves into a wall.


> so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote"

It's never the case that someone misunderstood what scientist wrote. Much like the scientific papers, news articles, including those reporting specifically on the discovery, have their own goals, and the paper being cited is used as evidence or argument for article's own "study". Except for press, the standard is rhetorical, not scientific, it's the conclusions and not the methods that are "pre-registered" at the start, and claims are defended by "hey it's just a point of view", not by statistical significance.

In your own example of worst offender: the scientific study was trying to establish and quantify the connection between X and Y. The summary article was trying to push the angle that "this institution is doing important work". It started with that conclusion, and the paper cited was just the first thing the author found that could be easily massaged into supporting that conclusions by rhetorical standards.

Same paper might get cited by journalist trying to push for "X is bad for you", and they'll do roughly the same as the summary article. And, same paper may be cited by someone claiming they have a miracle cure for Y, and they'll make a honest observation that "absence of X reducing Y is a common bullshit claim based on misunderstanding the paper [citation], that actually shows there's no correlation there, I mean look at the confidence intervals, even the author says that in text nobody bothers to read"... - citation may be honest, but the article itself is still using it to prop up a different flavor of bullshit.

TL;DR: don't believe news. It's bad for your mental and physical health (p<00.05).


> It's never the case that someone misunderstood what scientist wrote.

This is wrong. Reporters frequently don’t understand the science or the nuance in the science.

Reporting and science are two very different disciplines. Reporters rarely have a deep background in science and almost never have a background in the specific area that they’re reporting on.

Hell, even scientists have trouble accurately describing the work of a different scientific discipline.

Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.


> Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.

I'm not inventing them, but maybe conflating two sources:

1. Malice directly intending to hurt or defraud people. Probably not as common as how I make it seem.

2. Not caring. Well, I subscribe to the view that not expending effort to be accurate when talking to other person is as bad as slashing their tires (paraphrasing an old quip), so I very much consider bullshitting and picking a conclusion and then massaging facts to fit it, to be acting in bad faith too.


Nah, even in middle and high school it was plain that people sometimes misunderstood what the teachers meant. It's not like that goes away in adulthood; even when people "care", they still often misunderstand nuance or details, and sometimes even the bigger picture.

Honestly, it's plain weird to say that people never just make mistakes.

PS - worth adding that "I misunderstood" and "I didn't care enough" are not mutually exclusive. You can do both, so saying "they didn't misunderstand, they just didn't care" isn't a reasonable rebuttal. But even setting that aside, there'll be plenty of folks who care but still don't understand.


I'm gonna invoke Hanlon's handgun here: not attributing stupidity to what's adequately explained by systemic incentives promoting malice.

I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.


> I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.

The pattern of personal motives fits misunderstanding. You need to show that there is an organizational pattern of "malice" (your word, not mine), rather than an organizational pattern of "we are trying to publish quickly, and quality accidentally falls to the wayside". I.e., negligence, not malice.

You haven't provided even a shred of evidence suggesting there's malice at the journalist level. Every science journalist I have met genuinely cared about the science (which is why they were writing on it), but they didn't have time to learn enough about the subjects to understand they were oversimplifying things.


Not in science journalism, but I've personally encountered a case in normal journalism that I can only attribute to malice. It was many years ago, but it was so blatant I still remember it.

The 911 call went like this, according to its transcript. Caller: "This guy looks suspicious, like he's on drugs or something. It's raining and he's walking around looking into windows." 911 operator: "Can you describe him? What race is he?" Caller: "He looks black."

How the TV news reported it on the air was: Caller: "This guy looks suspicious ... He looks black."

Omitting excess verbiage is one thing. Omitting words that entirely change the context of the statement, making it look like the caller was racially prejudiced rather than responding to a specific question, is something else entirely. That was the last time I trusted reporting from that particular source (it was NBC, by the way).

My principle is that when someone lies to me, I stop trusting them. By lying I mean not just omitting details, or having an obvious bias, but deliberately telling me A when they clearly know that the truth is not-A. I could not see that report any other way but a deliberate lie, knowing the truth and attempting to make people believe the opposite.


this is such a bad faith Reddit-tier comment.

Try checking yourself on next few articles you read, and see how quickly you realize it's the actual reality.

Did just that with your comment. Seems like an exercise in discernment.

A lot of his takes on here are like this.

I have recently reviewed a paper that referenced my own article… but the conclusion was so off, I actually went back and re-read that entire article just to be sure there was no hint to the conclusion that the author derived. There was none. Uncanny experience.

This sounds a lot like knowledge graphs. Ideally, we can hyperlink to not only articles but concepts, facts, etc.

I’m a big fan of this approach.


Link rot has accelerated in the past decade. Links are great, when they work. If they are a few years old, they often don't. You certainly can't count on it. If you're referencing a paper or an article, it's probably better to cite the title of the article, the name of the journal or magazine or newspaper it was published in, the author, and the date. Then someone might have a chance of finding it again.

Wikipedia has solved this problem by using link + date for references and providingan archive.org link when the original is no longer available.

Doi url

That works for pieces that have one, although places that index those aren't always free.

Not just news stories. This is a huge problem with social media and forums in general too. Lots of people making bold claims about random things, lots of stories they say are based on a third party source, but very few actual links to said sources in question.

I still remember a recent example where one of those trivia accounts on Twitter posted an interesting story about some guy whose life completely changed after an accident, but neither linked to a source or named the person in question.

The only way I was able to verify it was true was through someone in the comments asking the platform's AI chatbot, and the chatbot providing context that I could research and verify...


> For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today.

I agree, this drives me crazy. Ironically, one of my favorite uses for Claude is to ask, "What study is this news article talking about?"

It's pretty good at digging up the source and related sources. And most of the time, if you read the source, the article is nonsense and gets everything wrong.


> There are still news sources doing this today.

It's the opposite, all big news websites do this. Fairly sure it's part of the policy.

I would say only small, niche websites link to sources.


> For a long time...a lot of online news stories would reference websites, papers, polls, etc. without linking to them.

The "$CITY_NAME Business Journal" websites are the absolute worst with this. They'll refer to something specific, for example "$BIGCO's 2025 10-K filing" and it will be a link. That link will go to the 10-K, right? Nope! It goes to another page at the same business journal. Maybe that page is a summary of the 10-K, but probably not. Maybe it's just the general index page for all the articles about $BIGCO at that journal. What it links to, it definitely won't be the specific thing described by the text of that link.


It isn't the transfer of information at all. What's actually happening is you're prompting experiences in a human instead of an AI using text.

Communication only works if you have multiple levels of representation and abstraction, including but not limited to - letter shapes, grammatical structures, style and register, stylometry, and subtext.

All of that is learned, and writers usually assume they can rely on that learning as the context for the text.

So you don't write to 'transfer information' like a network cable, you write to trigger experiences in the human version of latent space.

Factual information is one kind of experience. But even when that's the goal, there are always layers of implied relationship, social register, role, status, and other implications in everything that's written.

In normal communications the context - business emails, personal messages, mainstream journalism, fiction, and the rest - defines what acceptable language looks like.

The content fits inside that. But it has to fit the context, otherwise it lands in a semantic and psychological uncanny valley - like sending LinkedIn speak to a spouse on a wedding anniversary.

The real problem with LLM writing is that it's good at the technical layer - the grammar and spelling - and has some insights into the rest.

But the default content style is marketing and ad speak. And recently it's developed a weird and unique hybrid style which applies marketing fluff and pretension to technical content like code comments.

So you get one register instead of all of them. It can attempt others, but it's still too limited to generate them fluently. Sometimes the results are outstanding, but often it defaults to mechanical clichés.

So that's why it sucks and sounds so hollow.

Can it be fixed? Yes, but it's very hard work, most people don't have the skills, and it takes time - often too much time to be worth the effort.


I'm definitely being slightly too emphatic when I say it's "fundamentally the transfer of information", but I don't think that nuance is important.

When LLMs eventually get good at writing in the correct style for a given context, I'll admit that they have value in that way. But they aren't good at that yet. And even when they do get that good, I'll still dislike it for reasons that are more emotional than rational.


I don't care how good the LLM gets. If I know some text was written by LLM I'd much rather know the prompt - the seed of intent.

If the seed of intent is "convey XYZ details so they know them" then I can choose to go and learn those details any way I see fit - maybe even ask an LLM to summarise some data for me! - rather than having to ingest whatever their LLM use poops out and trying to digest the intent and content and figure it out.

It is about empowerment, rather than eating shit.


When sharing summaries from calls that I attended but my team has not but I think they would find interesting and I simply don't have time to handcraft it myself I use LLM and attach my prompt and the transcript.

Attaching the prompt initially threw some people..."wait, your admitting to using AI..."..."err yea, unlike you with that PowerPoint you sent me last week". I sense this is the right way to go imho.


This is one of the weird things about LLM use. It is a handy tool.

If you have found a prompt that gives a good result, then by sharing it with me I now have a tool that gives a good result.

The problem is not the use of LLMs. It is the "passing it off as your own work".


You're focusing on the wrong thing.

Output is not interchangeable with the prompt. In many cases, the prompt does not have the information the sender wanted to give you, and there is no guarantee that your LLM will give those information - or do it correctly - if you use the prompt yourself.

The entire value of here is that sender read the output and is vouching for it. This is where "bits of information" come from. If the sender cannot be trusted to verify and vouch for the LLM text they're sending to you, well, they're an asshole and you should rebuke them or find someone more considerate of others to talk with. Them giving you their prompt doesn't help you with anything.


I am focusing on the thing that is usually hidden. You are correct that for a full picture I'd need the information the LLM operates on.

But my point is that the prompt is the nearest encapsulation of the intent of the originator. Give me the data and the prompt. Vouch for the output if you like, but I want the source.


This is what's so amusing to me. People say they're not interested in the output of an LLM, only what a human has to say. But then when a human says "These words from the LLM are good, I vouch for them" the very same people say "If I wanted those words, I'd get them from the LLM myself, what's the point of this human at all?"

These humans are behaving worse than the LLMs at this point.

> If I know some text was written by LLM I'd much rather know the prompt - the seed of intent.

Here's the proximal prompt "Okay, take everything we've been talking about for 2 hours and apply those edits to the the final draft for publication."

What exactly does that give you?


If you give me that. Then you've given me nothing. Give me the same transcript the LLM has access to instead.

Anyone know of any good research articles on the Information-theoretical aspects of this? I find it a compelling topic.

While I agree with some of what you've said and the conclusion you've arrived at, in the end, I think you've missed part of the picture here with regard to "prompting experiences".

> Communication only works if you have multiple levels of representation and abstraction, including but not limited to - letter shapes, grammatical structures, style and register, stylometry

These are methods of encoding, there's no reason all of these can't be represented in an LLM from a technical point of view.

> and subtext.

This is the other half of the equation to me. Humans communicate by relating shared experiences, an LLM cannot have shared experiences. While it might be able to encode subtext that has been specifically called out and explained, it will never be able to encode the breadth of human subtext, especially that which is reliant on emotion.

I don't believe it is possible to change this until the point mankind truly develops a "wetware interface" to the digital world (and I personally don't want such a thing to exist).


That "marketing style" isn't a style, it's the lack of content itself. You just restated their original point despite trying to object to it, because it was correct and there is no way around that.

I don't agree.

Sometimes Claude's problem, such as when I ask it to summarize a long, complex session back to me, is it's too information dense. It uses weird invented terms to gloss over complex parts of the architecture instead of explaining them.

But no matter what - too dense or too sparse - it always sounds like Claude.


It misses the forest for the trees. It feels the need to highlight details not understanding what details are most relevant to a human reader and how to survey the larger problems in a cohesive way that emphasizes the right parts without cliche and undue emphasis.

It's never too dense. There are only more words but not more information.

If I describe all of the individual muscle contractions and joint motions required to walk across a room, it's hundreds of pages of data to say almost nothing. That is the exact opposite of dense.

By contrast a poet can deliver many concepts and many layers and even practically a fractal choose-your-own-adventure in only a dozen words. That is what dense is.


What already happens: - People give an LLM a bulleted list of points that they want expanded into a professional sounding document. - The receiver doesn't wanna read all that. They put the full document into an LLM and ask it to summarize it into succinct bullet points.

We've invented the opposite of lossless compression

Lossy expansion or bloat?

Depends on if you are buying tokens or being paid for tokens

Seven bullet points squeezed into five pages of text

Students where doing that for the ages when they have to answer questions like. "Answer the following question in not less than 4 pages"...

In that case it's supposed show that the student can write four pages and not just about transferring information.

A broken telephone.

TAM: $50T

It's been a huge boon to hardware vendors.

I first noticed this a year ago when I read a gmail AI summary, then glanced at the main text and saw it was AI generated. I feel like there's good fodder in there for a dystopian sci-fi story about a future where nobody communicates directly with one another, it's all AIs translating, but the AIs slowly start to drift.

>but the AIs slowly start to drift.

It's essentially the tower of babel. Each person will devolve to speak their own internal language only they understand. Each language will need to be encoded down to its meaning to be reinterpreted. None of us will know if the transformers are accurately decoding, or if the other person is accurately interpreting the decoding (which is arguably already a feature of human language without the computers in-between.)



Adrian Tchaikovsky has a good take on it as well, Human Resources.

It was in a gmail promotional video years ago already. Person A would use AI to turn their summary into email and person B would ask the AI to go back to the summary.

They knew it was going to be like that from the beginning.


Agreed re: the sci-fi story / trope.

I feel like a lot of this is a problem when someone technical is attempting to communicate a complicated technical subject to a less-technical audience.

I can only dumb a thing down so much before the description is useless (when you zoom out too much you lose the details). Even technical people who could understand it but are lazy / "in a hurry" use the summary, without thinking about what detail they are losing.

Even more infuriating is when they then reply to my email, having only read the AI summary, and ask a question that was already answered by my message.

This is the exact same thing that happened pre-AI, with the added step of wasting energy/resources on the AI summary in the middle.


I’ve been saying the same thing. I’m not saying all, but a significant amount of comms could be bullet points to the benefit of both sender and receiver.

Axios built a hefty business off this simple idea

This is exactly what I want: To communicate with me, have your LLM expand your message to include relevant parts of _your_ context, then I will have my LLM summarize it as briefly as possible with respect to _my_ context.

Just ask them for their prompt, and then put that in your own LLM with your own context. Why do you need their LLM's fluff?

It's worse than that. If it takes others longer for others to consume and understand what you're producing than it does for you to produce it, you'll never be able to communicate with someone efficiently. Communication breaks down at a fundamental level if if you can't keep up with the other side if outputting and they won't slow to allow you do do so.

I see this all the time now with LLM generated output. It's easy to have an LLM generate a chunk of content that can be dropped into a chat or comment, and when it took you 20 seconds to have something written up based on the shared understanding you and an LLM have about the context of the situation, but it takes other people 3-5 minutes to read and understand that content, that fundamentally doesn't scale. It's bad enough when one or two people are doing it, but if the whole team is doing it, the only way to keep up with the stream of information is to also consume it through an LLM. At that point you're likely to be missing much of the nuance, and the amount of errors will explode.

This can be alleviated by people reviewing the output of an LLM and making sure it both includes fundamental information that might be assumed by context and reducing it to the parts that are essential for the new context it's in. This takes time, but is extremely important.

Having an LLM write gobs of text to send to other people instead of doing it yourself is the equivalent of a low yield cognitive zip-bomb. Don't do it.


I don’t quite think this tracks. Perhaps you want to communicate 1000 bits that are well known and can be referenced with a 300 bit key. Then the LLM can easily retrieve the remaining information. It’s like sending someone a link to the Wikipedia page instead of explaining something yourself.

No, I don’t want to read LLM writing because it is BAD at it. It doesn’t really understand how humans think (because it thinks differently), and doesn’t seem to understand core principles very well (presumably due to the lack of world model), so it can’t write something humans enjoy yet.


That's kind of my point. If the 1000 bits are well known, then their inclusion isn't new semantic information. By pasting an LLM's output, you're deciding for your reader that they don't already know that information, and deciding that your LLM prompt is better than whatever they would to to obtain that information if they lack it. IMO, it's much better to give your readers the 300-bit key, and let them decide for themselves if they want/need to get more information, and if so, how.

If only that were the HN we comment on. If a post doesn't spoon feed an explanation for an initialism thats the tiniest bit off the beaten path (eg BGP), there's inevitability a comment about how would it kill the author to define BGP?

Mmm perhaps, but I think you normally have an idea of who your audience is. If this was actually how we talked then I would have just responded to this comment with “known-audience rebuttal. Example. Audience knowledge clarification. Example absurdity comment” and let you work it out. I think it’s normally known to you and the LLM and not known to the person you’re speaking to. Speaking like that would just cause confusion and misunderstanding.

> I think it’s normally known to you and the LLM

I disagree. Unless you gave additional information to the LLM yourself, the LLM doesn't know more than your audience does about what the meaning of such a comment would be. An LLM could certainly come up with something plausible, but it wouldn't necessarily be what you intended.


Yeah I don’t think it could decompress that answer, but most articles are explaining something and most explanations have been done before. The LLM would know in that case. I’m simply saying that you can communicate an idea worth any amount of bits with any smaller amount of bits provided you agree beforehand what they mean. LLMs have access to the entire internet, so we’ve had the chance to agree with them what every term means. But the person you’re writing to may have never heard of this, so you’ll have to communicate the full idea first before you can connect it to its small name.

Why not do both though? I often now write my human summary, then paste in also the content you could get by asking a bot (with markers for which is which). You’re perfectly welcome to ignore the bot text, or ask your own bot, but bots are pretty slow, so I also don’t want to wait for it to “decompress” that 300 bit blob to get the detailed page back. But my human summaries are also only intended to be 50 bits — compressed again to just pass the signal I mean to convey, and not the whole prompt needed to fetch the right info from the right place to substantiate the claim.

TLDR the length was the same curtesy of tl;dr before, just now with a different name.


I notice this in the attitudes of students towards reading. They will take a 10-page journal article and ask the AI for a summary and then read a summary that's the equivalent of maybe half a page. But if the article could have been half a page, why is it actually 10 pages? It's true that there is some boilerplate, but it's strange to me that people could think that 90% of what they're reading is (to use your phrase) "not true semantic information". It's like if you went to a restaurant and ordered a 10-oz steak and they brought you a little teeny bite of steak and said "Oh, other places will give you a bigger one, but most of that is just filler, we just took out all the superfluous parts." It's a worrying sign for our future if things like this are not tripping people's skeptic sensors and making them wonder if they might possibly be missing something.

A summary is about utility. They are definitely missing something but that doesn't mean what they are missing is useful to them at that time.

I mean I think that points to a related issue of people only focusing on a short-term notion of utility. The point of being a student is largely to learn things that may potentially be of utility to you in some way later, not just to do what meet your immediate needs (in the sense of passing the class).

I get what you're saying and I myself am guilty of surface level learning. But the idea that students should consume all the information is impractical. Learning what is acceptable to discard is part of being a student. I imagine very few students read every college text book cover to cover.

Beyond that people have different motivations and goals and only a limited time to achieve them. Basically I wouldn't be so quick to judge. Plenty of students have dropped out and gone on to do impressive things and that's a bit beyond reading only the abstract for a few assignments.


I’m not sure the 300-bit → 1,000-bit framing applies in all instances. The 300 bits may be a compressed cue to a much fuller idea. The AI can combine that cue with its prior knowledge to help reconstruct what the prompter was trying to express, with the prompter then verifying whether it’s right. Without the relevant prior knowledge for reconstruction, or the prompter for verification, it becomes much harder to know whether you’ve reconstructed the intended idea.

Unless that 700 bit was transferred on a separate occasion the inferred 700 bits is not true information, anyone could have reconstructed it from the 300 bits.

Not anyone, no. From an information theory standpoint, that it's possible at all to complete these 700 bits, only implies that anyone logically omniscient could. It's entirely possible that an LLM is capable enough to infer these 700 bits, and the human reader isn't.

Right but where is this information coming from, it can't come from the sender because an idea that can be conveyed in 300 bits cannot contain 1000 bits of information. You're just using the LLM to translate the idea into something more legible.

Alternatively those 700 bits are information the LLM added, but where is that information coming from? Is it noise? Random facts? Random lies? And who is the receiver even talking with if most of what they read is something the sender didn't know?


No, the 700 bits come from the sender verifying and vouching for the information before sending.

Unless that takes 700 attempts on average I don't think that actually works.

That assumes each bit is a coinflip, doesn't it?

Even Markov chain autocorrect tools do better than 50% odds*, and even GPT-2 was significantly better than that kind of autocorrect.

* at the word level; IDK how redundant/efficient language is when it comes to bits-worth-of-fact-claims-per-word. But "your cat is sitting on my" -> [mat, laundry, roof, head, belly, laptop, microwave, …] clearly has many bits of information, and a Markov chain will encode the most likely next word even if the user doesn't know what the most likely next word is. Verifying where the cat is sitting is also very easy, as is correction.


I assumed a coin flip, indeed in practice you would likely achieve far less than 1 bit per attempt.

Far more than 1 bit for most attempts: they need that just to be able to write coherent sentences, and a lot more to be coherent sentences on the right topic.

Some specific conclusions would be far less than 1 bit.

The average will depend on both the question and the AI.


This is about information from the sender to the recipient. It is fundamentally impossible to transfer more than one bit in one binary decision.

A LLM adds noise, not information. At least in this framing.


Frankly, it's a stupid framing.

LLM is not a random symbol generator (hint: training data is not random), and no reasonable person is going to just prompt an LLM and send its output without giving it at least cursory check (at the very least so that blatantly stupid hallucinations don't paint the sender as inconsiderate or incompetent).

That check alone can add bits to the final signal.


I'm not sure we're working in the same framing here. For one LLMs are random, sure you can fix the seed but you don't have to, you can even replace the RNG with true random noise.

So there are 2^300 possible ideas, only 2 outcomes from the cursory check, how do you get 2^700 outcomes? Most of those are just random variations the LLM added which is not a transfer of information. You would be lucky to even identify which of the 2^300 ideas was being conferred.


"Random" is too loose a word, but they were responding in a context where it meant coin flips.

LLMs are not even odds on all possible outputs, they are biased towards patterns which are upvoted by the training mechanism (at a minimum: the source material, RLHF, and synthetic data).

The information any trained model transfers to output, is information it gained during its training.

No single human is capable of having consumed all that training data.


There comes a point where someone isn't so much talking to the sender as having an unsolicited AI chat. Especially when most of the information didn't come from the sender in the first place.

It's actually hard to define the difference between information that comes from the model and just random variation, maybe something to do with the cross entropy between the sender and the model?


> There comes a point where someone isn't so much talking to the sender as having an unsolicited AI chat. Especially when most of the information didn't come from the sender in the first place.

Indeed.

The best case is a P vs NP situation: can the claims from the AI be easily verified, or not?

This does not excuse people too lazy (or overly impressed*) who fail to attempt the verification.

> It's actually hard to define the difference between information that comes from the model and just random variation, maybe something to do with the cross entropy between the sender and the model?

Mm.

Thanks to a philosophy course I did half a lifetime ago, I think there's a fundamental problem defining "information" in this context. It feels like it should mean "knowledge" because the discussions about Shannon entropy and transmission channels assumes there is an actual source-of-truth, but my conclusion from discussions about why "knowledge" can't just mean a "justified true belief" is thay I now don't believe we can do better than "belief"; an LLM can generate tokens that change your beliefs, but ultimately neither you nor I nor some annoying colleage who has made themselves redundant to the LLM, can be an oracle with definitely-true knowledge.

(I have of course tried asking an LLM about this thread; I don't feel it illuminated anything new for me, none of what it suggested made it into this comment).

* In the early days of LLMs, I was overly-impressed. Then I realised we were doing the same thing with LLMs today that we did with 3D graphics in the 90s, where every new engine was hailed as "photorealistic" only to be dismissed 6 months later when something better came along: https://archive.org/details/nextgen-issue-26

Only now it's every 11 weeks rather than 6 months.


> and no reasonable person is going to just prompt an LLM and send its output without giving it at least cursory check

I'm reminded of an old quote:

  The reasonable man adapts himself to the world: the unreasonable one persists in trying to adapt the world to himself. Therefore all progress depends on the unreasonable man.

Touché.

In information theory, each bit is a coin flip by definition

I realise I phrased this poorly.

I will try harder. Consider entropy.

The first sentence in this comment contains 32 characters; from the point of view of a naïve channel with no compression, that's 256 bits (given none require breaking out of the first bytes of UTF-8).

It did not take 2^256 attempts to construct the first sentence in this comment, because the generation process was not flipping coins per bit.

LLMs also do not emit bits chosen with a [0: 0.5, 1: 0.5] probability distribution.

From a compression point of view, the bits-transmitted-per-bits-in-message ratio can be reduced such that more likely messages use fewer bits than less likely messages. However, this requires the receiver to agree with the sender what the probability distribution over tokens is.

Intelligence is, amongst other things, a compression algorithm. If I can predict your next token, and we both know this, we can agree in advance that you don't need to actually send it.

No single human brain is able to predict the output of an LLM anything like well enough to do that.

In entropy terms: LLMs are noisy sources, their output does contain false statements, yet they add more bits of signal than of noise relative to a human alone.

Or at least, they can add more add more bits of signal than of noise relative to a human alone, but humans who blindly copy-paste the output of an LLM without checking are a pain and add zero value to whatever situation they happen to be in.

For some hypothetical scenario, writing software because I know they can do that, asking an LLM to write some code for you may easily give you 10 kilobits of positive information (code that mostly works), and -30 bits of noise (each bit being one binary decision's worth of incorrect choice by the LLM in what to write, i.e. bugs); if you as a user don't know how to handle the -30 noise that could easily be a totally useless app, but if you can filter out 30 bits of noise, either manually because those 30 bits happen to be your skill set, or even in some cases by prompting it again with the failure mode, then you get to benefit from the 10 kilobits of good stuff that you didn't have before.

In many (but not all) cases, LLMs can fix more than 1 bit of mistakes per follow-up prompt.


This is true - but also, models are absolutely terrible at writing articles and I don’t want to read them.

The issue isn’t that a 300 bit idea is padded with 15 KB of content. You can take any human-written article and reduce it by 90% with next to no information loss. What you lose is what makes the article a compelling read instead of a fact table.

I think the reality is that we will see quality long form AI-written content at some point. It doesn’t even feel like labs are particularly interested in chasing that now; code sells way more tokens. Right now the trend is that subsequent models degrade in writing quality as long as that pulls them up on coding benchmarks.


> You can take any human-written article and reduce it by 90% with next to no information loss.

You've unintentionally circled the error here. The "purpose" of an article extends beyond "convey this essential information".

By analogy, a textbook contains far more words than a spec sheet, but attempts to train the human to be able to easily interpret spec sheets. The so-called "information" content of both might be equivalent, yet one does a better job of teaching students.


Yup, it helps when the prompter reads and edits the AI-generated text, but usually they just skim it and send it unedited. Worse, it's usually not 300 bits -> 1000 bits (= a concise message covering all the relevant topics), it's more like 300 bits -> 3000 bits (= typical LLM verbal diarrhea hiding the relevant points in a wall of text).

I was going to write a post disagreeing with this on the basis of the fact that the reader lacks the background information the LLM has. For example, if I were to prompt "explain the proof of quadratic reciprocity using Gauss sums" most readers would need the entire LLM's answer (and much more, probably) and not just the prompt.

But then I realized that the reader can prompt the LLM with the same prompt for the same or equivalent expanded text. Most people don't do this as it's extra effort, but it's interesting to imagine a world where this is the default way of engagement with a text, assumed by both writers and readers alike.


> But then I realized that the reader can prompt the LLM with the same prompt for the same or equivalent expanded text.

That's basically what I've been asking my colleagues (so far a losing battle): Please don't send me AI-generated text. Send me your prompt instead. It is highly likely that I will understand it without needing an LLM, and if not, I can do it myself.


I like this example and it made me think, there is an analogy here to spec-driven development and vibe coding.

Rather than send 300+700 bits, like you said, send 300 (or less!) and let the human intelligence on the other side generate the result. Which supports the even older perspective: “If I had more time, I would have written a shorter letter.”

I’m not sure if this lands on anything very profound, but what about a pattern where, instead of codifying agent output at all, the only artifacts we share are the prompts. And the rewards (respect) accrue to those who generate the most generative among people and AI


As with everything, it depends how you use it.

I’ll often put a long stream of consciousness on the page, or jot down rough meeting minutes, then ask ChatGPT to “summarise this for an email”. The result is shorter, clearer and easier to read.

AI amplifies the habits of the person using it. If they’re lazy or dim, then it's like giving a monkey a gun.


> If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are.

LLMs do inference or computation among other things, so the remaining 700 bits can be something like that. The hidden implication in your claim is that computation adds no information content, which leads to an interesting philosophical discussion.

So for example, if I ask an LLM to give a proof or derive a new theorem from a set of axioms, according to your assumption, if it answers correctly, then I haven't learned anything new.

I am not really sure how to resolve this paradox in information theory.


If you input 300 bits into an LLM, and it outputs an additional 700bits of new information, I’d still rather you tell me those 1000bits than read ann llm’s output of 1000bits.

The primary reason is that human language is becoming a proof of work, that speaking out loud or writing directly indicates that the idea is important enough for a human to express. This is more costly than llm output, which is often just botspam.


Someone mentioned a similar thing elsewhere in the thread, that by communicating those extra 700 bits, the sender also implicitly vouches for them.

From a purely information-theory perspective, the simplest solution is to say that yes, any content derived from existing information carries no information itself.

From a realistic perspective in the context of people copy-pasting LLM output, my thoughts are that asking an LLM to research for you is more defensible, but it's still better to read the LLM's research results and write the important parts in your own words (partly because the LLM probably used way more words than necessary for the context).


> Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are.

Sure you can. LLM doesn't know what those 700 bits are, but you do. You may not realize it, and may not even know it at the time of prompting, but you do by the time you're sending.

Typical case is like this: you have 500 bits of semantic information to transfer. You give 300 of them to LLM, and get back the 500 bits you knew you have, and extra 500 you can quickly confirm are correct and relevant. Some of them are just dereferences of your input - where you recalled a pointer, but not what it pointed to. Some of it is information you never had before, but are able to easily validate.

You send that to me. I likely immediately realize the message was AI-assisted, but I trust you to be a decent human being, and not an asshole that lobs unverified LLM vomit over the fence for others to deal with. End result: you communicate 1000 bits of information to me, instead of planned 500, and you yourself learn extra 500 bits.

This is the optimistic scenario, but it does happen when LLM operator is not an asshole.

(Excuse the strong language, but I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself, so it's a topic close to my heart.)


I think you missunderstand. If the llm can "add back" 200 bits of information, these 200 bits are superfluous by definition. They can quiet literally be inferred from the starting 300 bits. And the llm has proved it.

If the receiver wanted these 200 added bits, she could infer them either herself or even use an llm to do it.


In principle, all of mathematics can be inferred from the axioms. The field of mathematics is not superfluous and something a receiver could infer themselves if they wanted it.

My brain does not contain all the information that can be added by an LLM; a human brain could contain it, even the biggest LLMs are about 1% of the (if you approximate synaptic count ~= parameters) parameter count of a human brain, but none actually will.

What my brain may actually contain is the information necessary to verify the (in this example) more than 200 bits the LLM claims to have added and trim out the parts which are false, retaining the (in this example) 200 "new" bits of new information added by the LLM.*

Concrete example: I am a software developer by training, though not a web developer. If someone who does not have any developer experience asks me to make a web app, I am forced to use an LLM as I do not know enough JS etc syntax to get it done myself. But as we all know, LLMs are only "ok" but not "good" at making software, so there are a lot of rough edges and outright mistakes. My experience as a software developer extends to detecting such failures and I can usually correct them.

The original person, someone who has no developer experience, can also prompt the LLM. Right now, this would result in something that retains all the errors, because they didn't have someone like me intermediating between them and the LLM.

I add bits by removing noise, the LLM adds bits but they contain noise.

I do not know for how long this will remain true, but today it is true.

* Feels like P versus NP to me. The answers AI generate are at their best when they're easy to verify. Then again, when they're easy to verify, they can be RLed to get good at this quickly and the need to verify goes down, leaving them still pretty bad at things that are hard to verify.


> If the llm can "add back" 200 bits of information, these 200 bits are superfluous by definition.

No, they're not. Getting those bits takes energy.

SOTA LLMs know way more than any individual on approximately anything there is to know (and what they don't, they can look up faster than people can). It's very easy for them to make the "missing" 200 bits explicit, rather than implicit, which in practical terms is the same as adding 200 bits that weren't there before.

Theoretically, an idealized omnipotent mind / AGI could derive the unifying theory from reading your HN comment on a phone screen. There is enough information there, if you were able to extract every bit of evidence available from it. But you are not. Neither am I. It would take us practically infinite work to try, solving this most cruel mathematical riddle.


> Typical case is like this (…)

> This is the optimistic scenario

So is it typical or optimistic?

> I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself

So why are you so eager to defend your fantastical scenario? It doesn’t matter how considerate you are, truth is the overwhelming majority of people aren’t and won’t be. We’re discussing reality here, not “what could be if we lived in a utopia which will never come to pass”.


> So is it typical or optimistic?

The optimistic case is "having 500 bits, giving LLM 300, getting back 1000, and learning extra 500 in the process". Real numbers are lower. People don't vouch thoroughly and don't catch all mistakes.

But reasonable people don't send every output from LLMs to others without giving it a cursory glance (obvious hallucinations or nonsense would paint the sender as incompetent or inconsiderate), and that alone eliminates the worst levels of noise. A cursory read and cutting out obvious bullshit before sending is enough to make the message carry more bits of information than the propmpt.

> the overwhelming majority of people aren’t and won’t be.

In my experience, the "overwhelming majority" are giving something between a cursory glance and cursory edit; whether the resulting message has more or less information than prompt then depends on how much noise LLM added on top. The inconsiderate people I deal with, they often send "net more bits than in prompt" outputs, but those outputs are also verbose and not fully filtered for bullshit, thus it's effortful to tease out the signal from noise.


How often do you modify the LLM output before sending it? If it’s less than 50% of the time, it means that by not modifying it, you have added at most one bit of information to what you originally wrote. (If you don’t understand why, think of it this way: Instead of sending the LLM response, you could send the prompt and one extra bit indicating whether the LLM response to the prompt should be modified, followed by the modifications.)

> How often do you modify the LLM output before sending it?

Me specifically, I never send anyone LLM output I haven't give at least a quick read (not skim, read) to make sure it's reasonable and there is no obvious bullshit there. And then I still mention it's LLM-sourced.

> If it’s less than 50% of the time, it means that by not modifying it, you have added at most one bit of information to what you originally wrote. (...) Instead of sending the LLM response, you could send the prompt and one extra bit indicating whether the LLM response to the prompt should be modified, followed by the modifications.

It's not the case, though. Prompts are not interchangeable with output. There is no guarantee that if you send a prompt, and recipient passes it to their LLM, they'll receive anything similar to what you did. It may have mistakes - different mistakes - or just spend focus differently.

The extra bits I claim LLMs can add to the message hinge strictly on you vouching for the response. Of course, you can just prompt an LLM, learn from the response, and then write your message clean, containing both the bits you originally had, and the bits you gained. But at that point, the LLM already gave you text containing all those bits - if you can vouch for it, you may as well copy it over and save yourself the trouble.


> It's not the case, though. Prompts are not interchangeable with output. There is no guarantee that if you send a prompt, and recipient passes it to their LLM, they'll receive anything similar to what you did. It may have mistakes - different mistakes - or just spend focus differently.

I’m not saying that the response is interchangeable, but that due to the data processing inequality, it cannot convey strictly more information than the prompt.

> The extra bits I claim LLMs can add to the message hinge strictly on you vouching for the response.

My argument is that if you vouch at least 50% of the time, the vouching only adds one bit of useful information – either you vouch or not.


> due to the data processing inequality, it cannot convey strictly more information than the prompt.

Only in the case where the LLM message is not reviewed before sending, and only if we assume reliable LLM (so that the receiver could recreate the same output if given the original prompt). This is not a realistic scenario.

> My argument is that if you vouch at least 50% of the time, the vouching only adds one bit of useful information – either you vouch or not.

The alternative to vouching isn't "not vouching", but "correcting and cutting out wrong bits and vouching for the rest", which means the single "vouched for it" adds all the bits that are in final message but weren't there in the prompt.


> Some of it is information you never had before, but are able to easily validate.

This is exactly the use case an LLM might (huge emphasis on might, depends on workflow, agentic vs. relying on contextual which can hallucinate) be good at and yet humans are notoriously bad at, because we are swayed by emotional responses and it is easy to have an emotional response to text that is programmed to look good for you and you alone.

> you communicate 1000 bits of information to me, instead of planned 500, and you yourself learn extra 500 bits.

Extremely optimistic. If this were the ideal scenario, you would USE the LLM to garner information ABOUT those 500 bits and then reframe them in a way that you yourself would put it. If there is insight, your "word" in your mental register now expands from the original 1000 bits to 1500 or 2000, and then are "processed" by your human brain that includes subconscious choices that are meaningful to the end result. There are tons of hidden semiotic data in your diction and wording (think resource forks in classic MacOS/HFS, only visible to the filesys) that is lost when you rely on another source to put together words for you; it's as if it is a game of Telephone. These are subtleties which you may intend for your recipient to receive and which are crucially important to your recipient and are irretrievable, it is intrinsically lossy. You have an alphabet soup of words, they cannot be put together by an LLM in exactly the way your brain did. We must rely on the fact that we ourselves put this together, the "aha" moment when an LLM does it for you is illusory and does not itself provide meaningfully important confirmation that you indeed say what you mean to say. Of course, humans say things and put things in way we do not intend to all the time. I still fundamentally believe this is more honest than relying on a third party that is not capable of understanding human emotional nuance to put together language for you, when language is and always has been a manner in which to dictate human emotional nuance.

> (Excuse the strong language, but I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself, so it's a topic close to my heart.)

Does not negate the fact that LLM output itself, at least when used to convey human emotions or thoughts, is lossy. A very highly compressed JPEG with added interpolation from an upscale algorithm might come up with cool details that were never present in the original, and may look cool to both you and the recipient but are not honest to the source material. You receiving 240p JPEGs on a day-to-day basis is irrelevant to this. For the purposes of communication, it is a massive error which has the potential to compound, regardless of whether or not you or the recipient believe this to be the case.


> If this were the ideal scenario, you would USE the LLM to garner information ABOUT those 500 bits and then reframe them in a way that you yourself would put it.

Yes, but at that point in practice we're getting into over-optimizing territory. In this optimistic case I presented, you could learn those 500 bits yourself and formulate a clean message yourself, with all 1000 bits in it, but since LLM already gave you the text, and you feel you vouch for, you may as well send it over and save yourself the effort.

In reality the numbers are probably lower, and writing the message yourself is IMO also a good way to be truly sure you vouch for the "extra" 500 bits, as it forces you to actually pay attention. There's a chance you'll find inconsistency in output, or in your own understanding. I don't begrudge people for eventually cutting the process off here, for practical reasons - it's the fuzzy line between accuracy and perfectionism.

> A very highly compressed JPEG with added interpolation from an upscale algorithm might come up with cool details that were never present in the original, and may look cool to both you and the recipient but are not honest to the source material.

Again, I think it's a wrong take. LLMs aren't pulling the information out of their asses, and you are also not able to express every information directly. LLM can "upscale" information and you can take a look and recognize, "yes, this is exactly as it was", even without being able to write out that "upscaled" version by yourself. Verification is often easier than direct recall.


The value of writing isn’t always to communicate new information. It’s often to align everyone’s assumptions. For instance, when I say casually to a colleague or an agent “this change will require a db migration” they understand it’s to my teams primary application database. If I submit a design doc to a company wide review which database is changing is critical information. If the 300 bits were truly enough your agent or junior engineer would implement the wrong thing correctly as they often do.

Clarifying which database you're referring to is exactly the sort of thing I'm talking about when I say "new semantic information". It should be communicated, and you can't trust an LLM to choose the correct database, you need to specify that yourself.

What I'm talking about is if you add "database foo" to your prompt, the LLM may then add text describing what that database is, where it is, etc. But that's not new information, it (hopefully) already exists in your team's public docs, slack convos, etc. You should just say "database foo" directly to your reader, and if they want to learn more about that database, they can do that themselves, or you can give pointers to them based on what you consider important.


> say "new semantic information". It should be communicated, and you can't trust an LLM to choose the correct database, you need to specify that yourself.

But I can most of the time, and correct it if it chooses the wrong thing. That's the whole reason LLMs are faster. Its the reason we can give a paragraph prompt and get a kLOC PR back but only need to correct about 5% of it.


>If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer.

A math teacher only has "this class is about math" to transfer, the rest is known. ;-)

Jokes aside, I think we're in a weird transition now where AI is used to generate text that looks good but is bad. In a few years people will know that and be more critical.

I think we went through a similar phase when DTP had it's breakthrough. Suddenly school papers were laser printed 300ppi times new roman and got more attention than better papers written by hand. But eventually that became the baseline.

I think Ai will make it so that well-written texts with clarity, good layout, correct illustrations, callouts etc become the norm, and will no longer impress anyone unless the information itself is actually good.

And if the information is good, it won't matter if it's AI generated or not.


Maybe things will change in the future, but I know where things stand right now. Would you rather learn about math by interacting with the teacher, by using the internet (including public LLMs), or by having the teacher tell an LLM "teach a math class" and then copy-paste its response to you? Option 1 is the best, but option 2 is at least better than option 3.

Oh it very much depends, your opinion there is certainly not always true and not shared by everyone; you only know where you stand right now with any certainty.

It depends on the teacher. I’ve had good and bad math classes, and LLMs today are quite a bit better than the bad ones. The worst human teacher in my memory didn’t offer interactions. He walked in, turned his back to the class, wrote equations on the board for 45 minutes, and then left. The best math teachers, the ones better than LLMs, are the ones who share the joy and sense of discovery and history of math, and not just the mechanics. But there aren’t that many teachers of that sort.

Learning by using the internet without LLMs is rarely very good, but more often than not in my experience sucks much worse than using LLMs. If you include using public LLMs in internet usage, then it’s not very different from just using LLMs that search the internet. I have heard that a huge swath of today’s high school and college kids are reaching for chatGPT before Google (which is incidentally OpenAI’s goal), and that many of them would rather talk to chatGPT than talk to a teacher. I’m going to refrain from making any claims, but I believe there are a lot of people who disagree with your ranking of the options.

Foreign language learning is one case where I love using LLMs, because it’s not typically an option otherwise. You can practice non-stop and have conversations with someone fluent in a language who will be infinitely patient with your mistakes. This is true of math and other subjects too; using LLMs to practice, so that the human teacher isn’t the bottleneck, to supplement and reinforce the human interactions, is usually better than using the internet without LLMs. The other reason many people prefer talking to LLMs is the lack of judgement. If you aren’t getting it and ask the teacher one too many basic questions, they treat you differently. Sometimes it’s necessary and helpful, and sometimes it’s harmful and takes a long time to change. LLMs don’t do that, they just explain and explain. That lack of judgement is a big reason many people prefer LLM interaction to human interaction.


Okay, sure. The difference between option 1 and option 2 is completely irrelevant to this discussion and I gave what I thought would be an uncontroversial take on it because I thought it didn't matter. The point is at least one of option 1 or 2 is almost always better than option 3. Option 3 is what I and others are arguing against.

Irrelevant? How so? And why did you present two different options if you thought it was irrelevant? I thought you were trying to make a point about human interaction. From my POV, your options 2 and 3 are much closer together than 1 & 2, so if you’re not suggesting that human interaction is preferable to LLM interaction, then what’s your justification for claiming any option is better than any other? I still completely disagree with your clarified point, and suspect many many other people do too, and are demonstrating that preference by using chatGPT for their classes instead of using either the teacher or Google. It’s a simple fact that interacting with an LLM can be preferable to interacting with some teachers some of the time, and LLMs currently often make interacting with the internet more pleasant and more efficient. Your ‘almost always’ still sounds to me like one person’s opinion that is not shared by all.

We're not talking about humans vs LLMs in general, we're talking about humans copy-pasting LLM output as if they wrote it themselves. It's less "LLMs are bad" and more "if you want to communicate something specific to me, LLMs are a bad substitute for your own writing".

Oh I see I’ve misunderstood. Your analogy is restricting the scenario to no interaction with LLMs. In that case, there are still some reasons LLM writing might be prefereable for students: for non-native speakers, and for math teachers new to match developing a new curriculum (sucks, I guess, but it certainly happens), just to name a couple.

One question to ask is why is your contrived LLM scenario any different than a math teacher using a textbook, or district/state worksheets? Math teachers typically do not write the course text or exercises, and never have. Very few math teachers do their own writing.


>Maybe things will change in the future, but I know where things stand right now

It's not an either-or though, LLM generated text already provides lots of value in many situations right now. It's also used for fluff, yes, but much of what humans write is fluff too, reporters often get paid by the word.

The point is that the value in a text has nothing to do with whether it was generated by an LLM or not. What matters is if it's useful or not.


While I acknowledge that this is meant to be framed in the context where the message is fully compressed, I think particularly when the medium is language (usually pretty compressable), then actually anti-aliasing the transfer of information is still 100% relevant! After all, AA is really about making better use of the information that you actually transmit. Compressed or otherwise, 1000 bits of aliasy garbage is not in any way the same as 1000 bits of perfectly antialiased signal.

I like to think of this idea of transfer of information (via words, let's say) from speaker/writer to listener/reader, as actually a fairly normal signal processing situation. One where the constructor of the message during it's 'rasterisation' from continuous thoughts to discrete words, can do the job well (ie: band-limit and 'antialias' the message, fully considering the target sampling domain), or badly (ignore the target domain and just speak/write from the source context, leading to 'aliasing' in the listener's received message).

I know this isn't how many people think about words, but I think the idea of antialiasing is relevant and something that people should think more about (including with respect to AI-generated text).


While I generally agree, the difference is I can also hand an LLM my pile of code and documents which is... a lot more information dense than basically anything I could write. Sure, a one-off prompt in a chat window isn't very helpful.

Summarization—especially of private context—is definitely one of the main exceptions to LLM writing being useless. However, it's still preferable to turn your private context into shared context and then give pointers to that, rather than having an LLM attempt to summarize it.

Oh fully agree. I hate reading LLM generated text, just wanted to provide a counterpoint to GP.

No, thanks. Depending on the type of information being transferred and just how uncertain the recipient's knowledge is, I may or may not need to add some amount of information to whatever I have to share. That information would be extra context you may need to have, on top of what new thing I am trying to explain. If I know I'm talking with someone who's all caught up, great, no need for context setting. Though even in those cases, there might be some terminology I want to define, or to provide a list of acronyms or whatever. That is boilerplate work, but is at least useful, if not needed. I know people tend to view this sort of comms as "technical" - they shouldn't be. So much misunderstanding happens because people are just not aligned on their shared context or medium. Of course, AI is not fully reliable to generate that context, you as a writer need to put in the work to review and align the result with your own understanding, or drive it thoroughly before.

All that being said, I acknowledge that people (me included) love to be sloppy in their comms (with or without AI) and then blame others for misunderstanding. It's also unlikely we'll change soon. What can you do.


Depends. If the 700 bits were arrived at by the LLM while spending a lot of tokens, and the result is "good", I may want it through you as a middleman because it used up your tokens and won't eat my subscription usage limit to ask the AI to supply those 700. If you spend the tokens and put the result online, plenty of people can spare their tokens because they don't have to ask the AI to derive it. Bonus if that result was run through some kind of testing and verification.

Obviously this doesn't really apply to super simple questions that the LLM can just spit out the answer to right away.


Sure, but then that's less "making an LLM write for you" and more "making an LLM research for you" which I don't think is what TFA is talking about

True. There is at least one more case: when LLM the other person is using has access to their context and information repositories that they don't want to share directly, and so the LLM text gives a peek at a slice of that, and that's not something I can recreate from thin air. In that case the LLM text may have utility for me.

Bits of information depend on the readers prior knowledge, those bits are not absolute numbers. That makes communication not just information exchange but also syncing of priors. So some amount of redundant information maybe needed/wanted.

So I spent a full workday investigating an issue with an agent and then I spent another hour prompting the agent to summarize the findings so that I can send them to a colleague. How does what you're saying come anywhere close to describing my work?

> you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer.

It doesn't actually follow, because maybe the LLM is smarter than the original writer (at least in the domain the writing is about) and hence really is able to complete the ideas in a way the writer can't. As an existing example, consider formulating a conjecture and having an LLM prove it. But I agree; if I wanted to read an LLM's output I'd simply ask it myself rather than read someone's supposedly-human writing.


Lets do a thought exercise:

You are stranded in a desert island. You start writing a message "Help, I am..." and pass at that point.

Somebody finds the message. They can no doubt come up with plausible continuations like "Help, I am Robinson Crusoe" or "Help, I am hungry" but they cannot create information. No matter how smart and how long you stare at the message, that is not going to tell you what the original person would have written.

Isn't it from Claude Shannon that information lowers uncertainty? Infinite regurgitation or massaging of data does not create new information. You will get the information form the LLM, not from that original person.


Once you've determined what the information-content of a message is, then you can apply information theory to it. But different receivers can derive a different amount of information from the same message.

Consider, for example, that if somebody doesn't know English at all, then before receiving the message, their best guess at what it is is some probability distribution over all English characters (or sounds, depending on what we assume them to know), and after knowing the first part is "Help, I am" that distribution might not change much at all. Therefore, they derived very little information from this message.

Going in the opposite direction: keeping fixed the knowledge someone starts with, there is an upper limit to how sure they could be (even if they are logically omniscient) in completing the message (that is, a lower limit on the entropy of their probability distribution) - this is what you're talking about in your example. But this limit only becomes important under these constraints - for example, knowing more about the person who wrote the message can let you predict it better, and if predictor A isn't logically omniscient, predictor B can do better than it with the same prior knowledge, just by being smarter than A.


What do you think about the prompt “The first 6 digits of pi are 3.1415”?

I think the prompt shows that you can't count to 6. :)

Also, "pi" is a unique constant name. You've conveyed far more than 10^[5|6] information in the subject phrase.


>It doesn't actually follow, because maybe the LLM is smarter than the original writer (at least in the domain the writing is about) and hence really is able to complete the ideas in a way the writer can't.

Then what's the point of the original writer?


I think you missed the import bit: "transfer of information from your brain to my brain." Including some third-party necessarily adds more that was not in your brain. So by using an LLM, you are transferring what is in your brain + some sludge that may or may not be useful for the other person's brain.

I agree. But there's still a valid use here for LLMs: If I have 1 Million bits of information and I want to communicate a synthesis of 500 bits to you, LLMs can be a viable helper. It is rarely done that way - I agree, but the way I use LLMs: Drop in all the relevant context information (PDFs, HTML, Markdown, Pictures etc. - up to usually 200-300k tokens), then compile a synthesis/summary prompt (usually 1-2 A4 pages of handwritten text. Then copy the output (1-2 A4 Pages), manually edit and send off (often, including an archive of the full original conversation, so the human on the other side can consult the unfiltered original prompt + inferencing, if wanted).

In other words: If I didn't reduce the 500 pages of text for you, you wouldn't know what I mean or what is relevant, or how to filter it yourself.


If only it were 300:700... it is more like 300:100000 and up. But then you use AI to summarize to hopefully retrieve the 300 bits. We should try doing this in a loop with a simple enough sentence to start with and see what pops out after every iteration. Fully automatic Chinese whispers.


"Chinese whispers" - a Chinese colleague teased me about using that phrase in a meeting, which made me question whether I needed to use that phrase in other contexts (e.g. conversation with family, in my home town, or with close friends).

I am not ready to use the Americanism, 'teh telephone game', since nobody knows what that means, yet everyone over a certain age knows what 'Chinese whispers' means.


Communication between people is much more complicated than that.

I get vague statements thrown at me with people expecting me to understand it.

Same with writing, setting up whole context to properly transfer 300 bits is always orders of magnitude bigger then just additional 700 bits.


I have 300 bits of information in rude direct form with some obscene language as well. I ask LLM to wrap it into a nice polite message. Resulted 1000 bits are sent in the email.

I bet no one would like to get direct rude "source" instead.


> I bet no one would like to get direct rude "source" instead.

You bet wrongly. Rudeness carries information.

“Fucking hell, how many times have I asked you to XYZ” is different from “G’day gov’nor, terribly sorry to bother you. May I remind you to XYZ? Would you mind doing so at your earliest convenience? My deepest regards, toodeloo”.

The former conveys urgency and annoyance while the former conveys that you can keep ignoring it (and straining the relationship).


Brit here.

The latter (with its twisted mix of Australian, Cockney and Kings English) carries a calm sarcastic tone which indicates ones displeasure far more than the former, more vulgar statement, could ever hope to achieve.

One is however, reminded that Americans simply don't get our sarcasm, frequently leading to some amusing cultural clashes.


> with its twisted mix of Australian, Cockney and Kings English

This is a great way of summarizing a lot of the TV we’re getting in the US that have British characters, and it’s annoying once you notice it.

See how the accent changes in Lie to Me between season 1 and season 3:

S1: https://youtu.be/bWyhsqh_e9s

S3: https://youtu.be/oPqOET_xCKw


gosh...this twisted indirect way of communication makes me nuts. Just be open and say what you mean and stop being offended so easily...

If you equate sarcasm with ease of offence, you are very much mistaken.

Regardless, in Blighty, such sarcasm is by no means twisted and indirect. It is as plain as the nose on ones face. As I said earlier, many Americans simply don't get such sarcasm, leading to comic clashes of culture.


Then you should push everyone to talk like a caveman. It really saves time and tokens.

But there is a reason no one is talking that way...


The latter one conveys urgency in the sense it burns more tokens on the governor's side.

The former also conveys that you are rude, which is not what you usually want.

I would much rather get the rude original message.

you just lost a bet

I don't think that necessarily holds, but it needs nuance. I tried to convey this within my company by giving a "guide" as to how to use LLM's for writing, across three modes:

1. Transliteration - roughly keeping the number of characters or bits, but translating to a different lingo, language or mental model (e.g. metaphors). Roughly the safest mode, but can still yield catastrophic results - it's safest if the author still provides taste and editing.

2. Compression - taking out redundancy to make the text more dense and more salient. The LLM chooses what to take out - and might take out the wrong things. More dangerous - but if you're happy with the salience and you believe the reader won't have time to read the uncompressed - it's probably safer than having the reader LLM compress without the benefit of your editing process.

3. Decompression - using the salience of your idea to add detail to the reader who wants to understand it fully, by utilising knowledge that is common to you and not common to the reader. This can be very powerful when there's no time to fully write the thing by a human - but it's the easiest to get wrong and to create slop. As an example - you could try explaining concept X + illustrate it through 3 examples. You know the examples are in public memory and easily retrievable - so you write your explanation of concept X, list the examples you want - and the LLM can take all of them, synthesise and bring the full package from your 300 bits to 1000 bits.

You are right that those are not the exact 1000 bits from the original brain, but they could contain 900 of the 1000 - which is still better communication efficiency than transferring 300.

I am however, more and more in the camp of fleshy brains writing everything, as my slop allergy rises.


This is a very good way of thinking about it! I like the nuance. Personally I don't usually include that much nuance when I write about this, because I don't expect my short comment to generate dozens of threads of discussion.

That's assumes a couple of things that are trivially not true:

- The assumption that both parties know about the same as an LLM does. An LLM know orders of magnitude more.

- The assumption that the output of the LLM is not refined over a few cycles.

The point is that you might give 300 bits of semantic information to an LLM, it fills it to a 1000 with perhaps 400 wrong bits. You correct it half a dozen times. It's now 950. You do the final touch ups. It's now at 1000. And it still took you 20% of the time to do it.


- An LLM has more knowledge, but it doesn't have information about what you specifically want to convey. Anything you give to the LLM may be good information. Anything the LLM adds is not additional real information, because whatever it adds can be inferred based on whatever you wrote. Or the LLM adds additional information that can't be inferred based on what you wrote, which is even worse because that's basically just misinformation.

- If you're giving additional prompts to the LLM to refine its output, then you're the one adding real information, not the LLM. The LLM is just rephrasing the information and adding noise.


- That's outside the scope of what I claim. The claim was that if you give an LLM 300-bits of information it cannot add more things that are relevant that the reader could not also do themselves. It certainly can. The LLM may add information you do not want to convey. Which is easily mediated by removing it or additional prompting.

- You are adding real information. The LLM is also adding real information. That's the entire point. It happens very often that an LLM suggest something to me that I did not know or simply did not think about. An LLM solved Navier-Stokes recently. That was most certainly not just adding noise. That's real information purely generated by an LLM. Information that was worth a million dollar price. Information that man centuries of mathematicians were not able to do generate.


Keep in mind that I'm using "information" in the information-theoretic sense. One could argue that because that particular solution to NS is provable, it is therefore implied by the propositions they started with and adds no new information.

And I'd still rather read a human's interpretation of the solution to NS than read whatever the LLM wrote.


> You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.

Exactly this. Just send me the prompt! ;)


Chatgpt, please write a response disagreeing with this post.

Be polite and professional, but assertive.

I’m not disagreeing with your preference for human authored writing, but I don’t know that I agree with the content of your argument.

If you consider that what humans are doing during conversation is a form of compressed encoding / decoding from some latent representation through a quantized signal then if you interpret it that as a compressed sensing problem you absolutely can infer to a very close approximation the original latent representation using far fewer than those 1k bits.


From an information-theoretic perspective, if I can effectively transfer a latent concept to you in 300 bits of quantized communication, then by definition the latent concept itself isn't more than 300 bits. Any bit over 300 in the communication I send to you is fluff.

Correct.

The real benefit is that the use of an LLM allows me to convey to it 1747 bits of scrambled information in an order that fits what's actually inside my head, and then have it unscramble that and convert it down to the 1000 you need. It can do that better and faster than I can.

It massively reduces the time it takes to make a short letter.


I think you need to take information compression into account. Compression naturally created by shared understanding of concepts and acronyms.

Can you explain to a child what it means when "GitHub is down" with less characters?


For a given information model, compression isn't a useful term, because "information" refers to a concept at its maximum compression. By definition, it can't be compressed further.

Take your point, I'm probably out of technical depth to throw these terms around on HN lightly.

What I was trying to say is that sometimes you need to provide additional context and information to those not using the same information model. Maybe it's not compression, but rather you have to include the 700kb of your information model with your message.


Information models aren't really something that one has, they're more of a semi-arbitrary frame of reference that we can choose. I'm assuming an information model where all public knowledge is accessible to everyone: you, me, your LLM, and my LLM. This means that conveying public knowledge doesn't convey any information (other than the fact that that knowledge is relevant to our conversation). This is a premise of my argument, and I'm fitting everything inside that frame of reference.

The LLMs don't have access to private information other than what you give it. If you give 300 bits of information to your LLM, the information it has is now "all public knowledge + 300 bits", and any text it generates is a subset of that information. By definition, it can't add any more information. If you give those 300 bits directly to me, I (and optionally my LLM if I want) now have "all public knowledge + 300 bits + my private thoughts", which is a superset of what your LLM has. Anything your LLM can infer can be inferred by me and my LLM. All your LLM can do is repackage that information into different text.

My opinion is that I don't get value out of that repackaging. I would rather read your packaging of those 300 bits rather than the LLM's packaging of those 300 bits.

For this current discussion, we could use an information model saying that your LLM has a knowledge base private to just you and it (e.g. local files or other conversations), and say that's separate from the prompt you give it. It sounds like maybe that's the information model you're thinking within.

In that frame of reference, maybe your shared knowledge base has 200 bits, and you type 100 bits into your LLM's prompt. My argument would still be that you're "giving" 300 bits to the LLM, and you should instead give them to me by sharing your knowledge base, or giving me the relevant information in your knowledge base in your own words. The latter option is definitely more work for you, and is the weakest point in my argument, but I'll still hold that preference.


> GitHub is down

"No work today"


I love this example, partially because it jives with my conviction that LLMs are the ultimate translation machine. Ever since the embedding model days, it is clear that these models are amazing at representing meaning as math. The fact that LLM's most salient use is for coding somewhat agrees with that. After all, what is a programming language but another language? We instruct people with words and machines with code.

I like your premise but I will point out that communication isn’t just about conveying information.

It is often about persuasion and that sometimes benefits from framing effectively, which I think an LLM can help with given the key points you’ve got.


Using an LLM to write something persuasive is a good way to persuade me that you don't care about the topic personally, and therefore I probably shouldn't care what you or your LLM think.

Saying the hard part out loud: we give more value to the people with 1000 bit brains than we do to the 300 bit people. There’s a strong incentive to write like you’re the former and not the latter. Freely available LLMs give people the means to match the incentive.

What if the bits I'm transferring are "I don't remember the precise syntax, but here's a link to the relevant documentation."


I love it :)

If you have 1000 bits of semantic information that you want to transfer but your default communication combines it with 10,000 bits of noise. Giving it all to an LLM and iterate on reducing that noise while making sure the 1000 bits is still present would enable you to communicate more effectively.

Overall, ideas are ideas. I'm not overly concerned with the fact that it was you who had the idea, as long as the idea is interesting. I don't know most of the people who write the things I read, so it seems to be of no consequence to me at all if they wrote it, as long as it is interesting. LLMs are notorious at creating things that are bland and vacuous, but they by no means have a monopoly on it.

Be the source human, machine, or dolphin, if they write a good article, I'm prepared to read it.


LLMs are wonderful at adding noise and okay at removing noise. My point is that they're not very useful at adding signal. If you wrote 10000 bits of noise and 1000 bits of signal, I would rather receive those 11000 bits from you, and if necessary ask an LLM to remove noise based on what I consider noise. If you can point out to the LLM what it should consider signal vs noise regardless of context, it should have been easy to not write that noise in the first place.

I think there are a lot of situations unaccounted for here

- people writing in a non native language

- people insecure in their writing

- people not used to writing in industry terms

- people with the curse of knowledge that are aware that they can’t write for a general audience well

Surely others too. None of those mean you have to read it, but I have gotten immense value from reading some things people have had ai write (and I’ve seen a ton of junk as well)


> people writing in a non native language This category of people are going to be most betrayed by LLM output: 1) the receiver loses the signal that the sender might need to be queried to find their real intent and 2) the sender doesn't have the ability to determine if what they are sending is what they mean

> people insecure in their writing These people can grow up, I don't care. Not a good enough reason to send a slop grenade.

>people not used to writing in industry terms Similar to the non-native speakers, but slightly less in magnitude. They can educate themselves though.

- people with the curse of knowledge that are aware that they can’t write for a general audience well These people probably can get some value out of it but they should take care

Still not really a good enough reason in the end


That’s like, just your opinion man.

I disagree, but I think neither of us are going to benefit from continuing this discussion.


isn't it more like: i have 1000 bits, i transfer 1000 bits but depending on the person, it might be lossy, so they only understand about 700. they then come up with the 300+- on their own, potentially putting them over 1000 or they come back and ask questions to fill in the blank. the bits don't ever have to be bit identical.

They're lossy. I might make an esoteric reference that no one gets, or an analogy that doesn't quite land, and then have the LLM help me come up with something a bit more understandable to a general audience. There's a difference between dumping the output that you spent 5 minutes with, and taking an hour to craft that perfect analogy.

Sure, but that seems orthogonal to what I'm saying, no? We could continue clarifying the definitions of information in this context, but my fundamental argument is LLMs don't add information to my writing that couldn't be added by the reader themself.

This of course has the potential to change with personal LLMs that can have shared private context with me. However, that isn't a defense for sending people AI slop, it just turns it from "LLMs don't add value" to "LLMs may add value when used judiciously."


you think like such a techie. consciousness is not analogous to IO. if writing were only the transference of something from point A to B then what is the technical function of poetry, a question (the open ended kind), a pondering, a wondering, and so on? Furthermore, language is lossy and introduces a large degree of subjectivity, mystery, and uncertainty nomatter which words you choose. Words are by nature lower res/on a lower ontological domain than thought. So your premise is preposterous on both a practical and theoretical level.

Furthermore not all writing is for another to consume; nor even for the author themselves to consume. That is to say it has meaning ipso facto, not dependent on transference, as ritual.


Art is an abstract method of communication. When you choose the words of a poem, you're (hopefully) doing it to convey a feeling within you to the reader. If you write a poem about a beautiful spring day, it's probably because you experienced one, or you're remembering one, or someone was telling you about one, and that evokes feelings within you that you want to put into words, right? Surely you wouldn't write a poem about something you don't care about in any way?

When you talk about something you're wondering about, you're saying that you're missing information. Your ponderings are dancing around the void in your knowledge, defining its boundaries, and maybe imagining what answers might be able to fill that void.

When you put your thoughts into words, they're insufficient. You have so many ideas swirling around in your head, and you can never put them all on a page in the fidelity at which they exist internally. But words are the best we have. Whatever words you write are your best attempt to convey your thoughts to me (barring other media). You're distilling your inner voice that speaks a language only you can understand, into an outer voice that others can understand.

I don't think I'm exactly refuting you here. I think what you've written makes sense, and caused me to think about many things, more so than any other reply to me today. But I also don't think your comment is refuting the point I was trying to make, mainly that LLMs rarely add value in human-to-human communication.

I could probably have pasted my comment and yours into an LLM, and it would have come up with a clearer thread connecting my words to yours. But that thread probably wouldn't have been any of the ones either of us saw, would it?

Thanks for adding a new perspective to the conversation :)


To zero in on your last question, I do think a very good LLM can surface the unconscious (even completely unstated or hitherto nonexistent) "connecting thread" linking two subjective thoughts and make them explicitly known. And so what if they were not something either of us had in mind?

Surely the ability to do that is worth taking note of.

I am not advocating for letting LLM's write for you, to be clear. Sentiment wise, I largely agree with you. Just not with your total writing off of the possibility that it could serve.

It's easy to imagine an LLM aiding the communication between a mentally disabled person and their parent/caretaker.

Or, perhaps, some day, between animal and man. Who cares if the mediating component "hallucinates" some particulars of expression if it achieves the goals both want, which were previously impossible?


If you read anything you're reading superfluous bits.

How do you explain professional speech writing?

Professional speeches are usually equivalent to LLM slop in terms of information content. They're written to sound nice first and foremost, and communicating information is a secondary goal.

So... We should only communicate in compressed formats?

Yes, but specifically compress the concepts, not necessarily the text encoding. "100 commas" and ",,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,," may be different amounts of ascii data, but they both represent the same concept that's worth around 20-30 bits of semantic information.

Just use LLM to filter out those extra 700 bits, and you'll happily end receiving 100 bits or original information and 200 bits of slop.

genius

This is a great analogy, thanks.

That means folks using LLMs starting off with 300 bits KNOW that they lack the full payload of information to transfer to you. IOW, they know they need to transfer much more than 300, so they use LLM to fill those gaps. That's the crux of the slop universe out there. Folks are using LLMs for the 700 bits on top of their 300 bits and passing off the full 1000 bits as their own.

I echo the writer's sentiment. "I don't want to read the clanker's 700 bits. I want only your synthesis." (I can get the clanker to generate those 700 myself. Unless ... unless this whole LLM slop market is all about saving you the time to get an LLM to generate those 700!)


I will give you a use case where this is absolutely not the case.

I have a bunch of CLI utils I run for various clients and their peculiar setups. They now have man pages with descriptions and examples in them because the LLM went and read my code and did the needful.

I no longer have to re read my own code, rather I can just use the manual page.

Format and description came from semantics and context that (barely) existed elsewhere and I was not going to retain or transmit, but I have now.


But then the LLMs aren't adding new information, they're just reading your code and translating that information from e.g. python to english. Rather than sending someone an LLM-generated doc to someone, send the code and your own personal thoughts on the code. If your reader doesn't want to read and understand your code, they can ask an LLM to analyze it, within the context of their specific use case and your personal thoughts if any.

> e.g. python to English.

You're making a big assumption that the code is what is being executed, and not a compiled binary.

Where is the code: My repo? the clients? If it's in mine, the client does not have access and the CLI is a first stop to debugging. They arent in the context of written docs, more likely a production error from a log (thats now spitting out a message to check the CLI).

Less steps, less tools, more context in line and available in an interface your already using.

> they can ask an LLM to analyze it, within the context of their specific use case and your personal thoughts if any.

Or I can skim the man page it generated and make sure it looks good. The "work" (the tokens) dont have get spent over and over again.


I never liked information theory because information theory as Shannon envisioned it fundamentally did not deal with semantics.

AIT tried solving it? But AFAIK it's a lot of pretty results with not much real application.

A better approximation is something of a "shared model"; then you can actually state things like, the transfer of information sometimes is "trivial" because, well, it's right there in your compressor/decompressor.


I was just trying to understand this about a month ago and it is interesting how little there is in terms of semantic information vs Shannon.

An Outline of a Theory of Semantic Information by Carnap was the early attempt.

Fred Dretske wrote Knowledge and the Flow of Information in 1981.

Luciano Floridi has a few recent books.

I couldn't find much else. I don't think AIC really solves the problem of meaning either.

I think the Dretske book was the first time I really understood where Shannon was coming from but I gave up when it got to his actual semantic ideas.

I think I ran across a recent paper that motivated trying to back track what work had been done in this area but I don't recall the name of the paper.

I've have shelved all this for now as over my head.


My understanding is that to unambiguously quantify information, you need to have a known model within which that information fits. In this context, talking about LLMs adding value (or not) via inserting new semantic information, we're assuming that public knowledge on the internet is not new semantic information, and the quantity of information of contained in text talking about public knowledge is equal to the ~32 bits needed to point to that knowledge.

"Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700,"

You absolutely can if that information is in the code, which it often is.

There should not be that much in the code that needs further elucidation.

Some stuff definitely - but not much.

Usually you need the code and architectural summary + that stuff.

The AI is not very good at it but it will get better.

I think the debate here is about a few different things.


That would be true if there is only one receiver and if i could exactly predict how the receiver would interpret these 300 bits. But i could use LLM to expand these 300 bits to 1000 bits of more explicit information that gives enough redundancy and context that minimizes risk of misunderstanding and increase easiness of understanding, then validate such expanded message (and possibly do iterative corrections if LLM did not understand those 300 bits correctly), and finally release expanded message to (potentially large and heterogenous) group of receivers.

If there's a significant risk of misinterpretation when writing only 300 bits, then the idea you're trying to convey is worth more than 300 bits. In the process of prompting the LLM to add information, you're giving information to the LLM that you could instead just give to your readers directly. If you're just using an LLM to expand your thoughts in the hope that you get a better piece of writing, maybe you should put more effort into your own writing.

(This is all under an information model that assumes the LLM and your readers have equal access to knowledge, which I probably should have made more explicit in my original comment.)


Saying something for years doesn't make it right. Has nobody ever challenged you on that? Our brains prefer to read enjoyable bits, not raw information. If you only got those 300 bits, you aren't going to bother putting them into your LLM to generate 700 bits to make it a more enjoyable read, you're just going to struggle through the 300. Maybe when browsers come with built-in automatic text puff-uppers, then you'll have a case but almost nobody does that now.

Here's a clearer example - would you rather learn a concept from a research paper or a textbook or blog? You say the research paper but they're dense and hard to wade through where-as blogs and textbooks are more wordy but hold your hand, which is something that helps humans learn.


Plenty of people have challenged me on this! There are a lot of exceptions we can come up with (summarization and private context are the biggest examples), and everything's changing in various ways as AIs get more capable. Usually, conversations about people pasting LLM output focus on things like AI generated blog posts, people who have no taste in writing and want an AI to rewrite everything for them, people who are just plain lazy, etc.

> would you rather learn a concept from a research paper or a textbook or blog?

I pretty much always read blogs first, and then move to a research paper only if I want more details or care enough about the subject to verify with the original source. Typically this is because research papers have too much information to be approachable.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: