I still have this naive notion that we don't need LLMs for code generation and editing. Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
If an organization is paying $2400/year per developer for tokens and a highly intelligent editor/IDE comes around that charges $1000/yr and gets more output at a fixed cost, its a no-brainer of a decision.
You know what, considering there was a recent "small" open weights LLM released recently that meets 90% of my coding needs I'm inclined to agree.
Qwen3.8-Flash-Next - relatively small, it runs on 6 6 year old GPUs on my home PC happily running 5 simultaneous 262k sessions with additional 10 cached in RAM (bought back when you didn't have to remortgage your house for Ram) and it has been the first local model that is not a toy.
But there is a class of problems where I still reach for Anthropic's fable...
However, I have a hunch bordering with certainty Anthropic is achieving such great results by doing a lot of harness tricks.
For example opus 4.8, is not much better on coding than before mentioned Qwen model, but gets amazing results on factual knowledge stuff (the knowing all works of Shakespeare thing). How hard would it be to add a general knowledge RAG to requests that contain relevant questions and beat all benchmarks like that? Not very hard.
So I think there is big innovation to be had in harnesses, routers, inference and so on.
As to money spent on AI per developer my current client (a fortune 200 software company) spends $500 per month. That is $6k a year. A lot more than your examples. And many people run out of their quota pretty quickly.
> "Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?"
If the implementation brief says "attempting a reconnect in this handler would be a wild goose chase", the model needs to know enough Shakespeare, at least indirectly, to understand that expression...
Forgive my naive understanding of LLMs - but how do you get semantic understanding of a codebase, such that it knows what changes to make/why/where, without a wider understanding of language more broadly?
I'm using the word understanding loosely there, but I couldn't think of another word.
Looking at what Jev has shown, and what is being done in the space. I think you are right that there’s probably a lot of room for performance improvement on specific tasks and workflows that will be happening in the next few years
I also imagine it could be a big shakeup if all of a sudden models could run on CPU. Imagine running an Astra-level coding agent, locally on your laptop. All of a sudden GPUs wouldn’t look as valuable, if you don’t need them as much
We are still some time away from that, but it seems like progress is being made
> If an organization is paying $2400/year per developer for tokens
My employer is spending $50,000/yr per employee on tokens, and they're not alone
We’re conditioned to interact with language models as chatbots and in that sense strong language understanding (implicit - some sort of world knowledge), is probably necessary for that?
But I’m sure we can have a small model that’s really strong at programming concepts, JavaScript syntax, and that’s about it. You’d interact with it differently, at specific seams in your code base - review a PR, merge two functions together, investigate these logs.
Or maybe I’m just not adequately absorbing the bitter lesson. Idk
I think the amount of knowledge to correctly work on code is more than you'd think, because at the end of the day writing code without an understanding of the environment it exists in/for is likely to not fit the problem correctly. Maybe it doesn't need knowledge of _Shakespeare_ per se, but if you were working on a virtual tabletop having knowledge of tabletop games can help with identifying the right implementation to use, knowing what kind of constraints to consider, etc.
CodeRabbit emits LLM generated poetry in its comments so maybe more than you'd think?
This is exactly what the frontier models are trying to do!
Check this graph:
I have a feeling that when all the AI-hype dust settles, what you describe will be the killer app of AI. That, ChatBots and unstructured data processing. Huge productivity improvements but not the sci-fi hype of today.
I don't need AI in the same way that I don't need autocomplete. I can definitely program without autocompletion, but I'm a lot slower than others who use it
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
What language do you plan on prompting it in?
It is very difficult to predict ahead of time what knowledge a model would need to understand a prompt to generate a program. It could refer to all kinds of real world knowledge referring to the kinds of entities you want the program to model.
I think this already happens through Mixture of Experts which is now build in to ost models.
But finding out what an LLM needs to understand from a business side to write your code good, is an otpimzation which no one cares currently.
I'm pretty sure we either stay on big full frontier models for a long time, just use them for everything or we will start to see more and more people doing finetuning/project specific training like java + german + english + business contxt xy;
It will be an indicator for the whole industry.
I suppose how much literary support you need in your model depends directly on how erudite the comments are.
Would there be problem domains in which the more educated LLM would perform better? Are your names directly related to concepts from said domain,
LLM comments: "I think it may be a potential bug that the sum VATAddedTax gets added to the TaxFreeItems".
I mean it seems a bit unnecessary but also maybe it can help in unexpected ways.
Because we don't need code
func randomName () { desired machine physics }
Everything around "desired machine physics" is superfluous wank; historically a biz case stored as code when some UI could feed biz case params go a function generator
Come on we know what we use computers for; media consumption and 2D data entry/review. Locally we just need a core engine for geometric transforms of visual state. What all these languages give us ability to create such a generic VM filled with customized semantics that mean nothing to solving the problem but plenty to a clever coder.
Kind of like Unicode we need distilled geometry primitives like "teapot for text" and desktop metaphors and to let people put the superfluous wank at the presentation layer
Which text used to be so making UI out of layers of text, OOP, and such made sense for decades
But we're just engaged in bloating system state through def jargon_to_encapsulate { desired machine physics } when we already know it's going to be simulated 3D or 2D visual transforms. We don't need to capture all those states in code verbatim.
Things like Jev are the future of models. Fine tuned on transforms given a context. "So you want to replicate GTA5? Here's a data set of geometric shapes and gradients constraints from all observed xyz" pipe that into your local renderer
We're entering the phase of software engineering (and engineering generally) where we realized we been dramatically over playing the song and can strip out entire asides and digressions, circumlocutions of provenance, to tighten up pacing and improve enjoyment of the outcomes. Hopefully. Or we kill ourselves. Through social squabbles (political, economic, religious, whatever) due to laziness to learn etiquette, and environmental destruction.
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Probably not unless you're writing tooling relevant to literature or prose, but I can't imagine trusting jetbrains (or any ai studio) to curate this.
I'm not an expert, but I think that:
1. Storing Shakespeare's work costs almost no $ in regards to disk space.
2. If the prompt doesn't include "Shakespeare" or relevant terms then no regression is performed for that topic and therefore there is no effective token cost.
Someone may correct me, but I think it's not a big $ win to exclude relevant topics from the models' overall capabilities. Instead you'd tune weights so that #2 better identifies what is or isn't among the relevant terms on which to run regressions.
> Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
IDEA already have small LLM for one line code completion IIRC.
But the gain people want from LLM is generally "here, add this entire feature" or "here, go thru every dependency's changelog and update code to work with latest version". Those are not small LLM tasks
I have hope we will get there eventually, once all the hype/wealth extraction/boys club giving all their buddies money cycles end, and the specialized tools with real value start to emerge.
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
That's needed to interpret the 10000 monkeys typing requirements in the various corporate product roles /s
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Turns out that actually - no. Researchers have managed to prune half the Experts in a MoE model that had a low probability of getting activated during coding tasks, resulting in a more focused model:
https://arxiv.org/abs/2607.16721
Main benefit is that it greatly reduces the amount of RAM required to run these models. Of course you could just cache those unused experts on disk instead, but the main point here is that you know which ones matter.
But aside from that recent models, like Qwen3.8-27b are reportedly more durable under heavy quantisation, e.g. 3bits or even ternary. With additional techniques like TurboQuant, you can feasibly run these models on consumer hardware - even if at 1/4th the speed you'd get from rented infrastructure.
VS Code has extensions such as Kilo Code or llama-vscode which let you work with local models much like you would with cloud based solutions.