In short, the more deterministic, automated, checks the better. AI can deal with a pedantic language. I intend to add statically verified structured concurrency, units of measure, contracts, and eventually more and more formal methods into the language so it can be a familiar TYpeScript-like base with as many static guarantees as we can fit in.
I also think that fine-grained isolation, which Zena gets via Web Assembly, is critical for limiting the capabilities of generated code and the blast radius of bugs, vulnerabilities, and non-aligned behavior.
I do have an optimistic hope that a language also optimized for humans, readability and simple semantics especially, has value in the future, even when most code is generated. We'll see about that.
I love this. I was thinking about a "cleaned up" typescript for a while now, and this seems to be it. I believe this can work better as an "ai-first" language than some other attempts I've seen that try to reinvent the language from scratch.
One thing I would love to have as a feature is native compilation.
Ironically, since AI can build an STL and tooling (up to a full OS!), I think now presents an opportunity for a language that does start from scratch, at least targeting hobbyists (besides those specifically looking for novel languages).
Why not languages that are designed in that way, like Ada. Or if one wants less of rigidness then a static typed language that makes it hard to shoot in your foot like Rust,Java, C#?
JS derivatives are that, a derivative to a scripting language.
Agents are fairly good even at languages designed to trick them. I built one[1] and it does make for a good benchmark[2] to see which LLMs are actually good. I think that a language which is largely similar to others will be a piece of cake for most and any advantage that an existing language will have will be minor enough to not matter.
Btw if folks have ideas of how to make Killswitch even harder for LLMs I’d appreciate them.
From experience with Zena, this is not true at all. Opus, Fable, Gemini Flash and Pro all barely make any syntax mistakes after a little is in context, and those are caught extremely early.
The one thing I do see sometimes is that agents sometimes don't take advantage of added features, but that's partially because the Zena code base doesn't use them as much yet. I'm working on skills and linter-based suggestions to use better patterns.
You should have working programs that demonstrate the features you want them to use, and then the skills. Working programs they can mutate in an RL gym.
This is absolutely not the case in my experience. I am building a very large embedded domain specific language for describing distributed systems. It looks like a small subset of Elixir, but with object-oriented syntax in a lot of places. (It’s called a choreography; there exist many other choreographic programming languages.)
Even though this programming language is absolutely nowhere in any large language model’s training set, they have so far done extremely well at extrapolating from the small set of examples I’ve given it when I need an agent to generate some tests or whatever for me.
I think starting with a familar typescript-like base language is a good approach to this. This should be familiar enough for LLMs for the most part as long as additional features can be explained in a succinct system promopt/skill.
> bc agents will naturally be bad at it due to a lack of examples.
Can we please as a community stop parroting these false premises as a basis of every argument against doing anything new? It's plainly obvious to anybody that uses LLMs on a regular basis that it's not true.
I’d love for language environments to support encouraging LLMs to specify more when they write code. Why should they first write a complex function or a class and later bolt on a test?
I’d love for a class in this future language to come packaged with tests, invariants, fuzzer parameters, profile targets / performance budgets with realistic inputs (on this 100 element array this should take no more than X clock cycles), race condition stress tests etc.
The compiler (or even linter) should optionally run some / all of these checks and succinctly report back (with knobs so the LLM can manage wall clock time).
Adding each of these should not be follow on steps.
Beyond this, debug hooks should be trivial to set (in code itself), so the LLM can trivially say show me the stack after the 9th time this function is called on this input to the program.
I think I get your point but I think it's reductionist to the point of being incorrect. LLMs must be better at some semantics than others. Programming languages don't have random semantics, they have what matches the world and what matches our languages and so on. And the current frontier LLMs aren't so generic that they can predict any phrase no matter the quality of the content and grammar. Concretely I mean some languages are harder to reason about (predict) than others.
LLMs are good at optimizing towards local goals. Getting types right at compile time is a local goal. Entry and exit assertions are local goals. Unit tests are local goals. So those constructs all help AI-generated code.
Matching a desired output is a global goal, but even that sometimes works now.
Someone sent me a LLM-generated JPEG 2000 decoder. They got Fable to generate a decoder that uses a GPU to get the same answer as the reference implementation gets on the GPU.
LLMs are generally more tolerant of tedium than humans, but they make mistakes more often on repetitive mechanical tasks. I've had Claude write PTX directly once, and Claude just wrote majority of it and commented something along the line of "repeat this block 7 more time with these minor changes" instead of writing them out, so the code didn't work.
So, no, replacing compilers with LLMs is probably a worse option than having them code a compiler/programming language.
Oct is the first programming language that I made with Codex. It started out as "Octave Modern" and was never intended to be a language for LLMs in the first place, but rather a teaching language that I've been thinking about for a decade because of my frustration with academic code and specifically reproducibility, with Python and Matlab in particular.
But it turns out the same design choices that was made to prevent bad patterns from academic code also made it pretty good for LLM coding: statically typed, immutable by default, GC'd with fast compile and runtime because it compiles to Go, along with features designed for scientific compute like SI units and builtin graphing etc.
But now it just took a life of its own, so it has extra features like templates/concepts, iterators, async/await, database query, build system for C/C++, SystemVerilog/WASM(WIP) codegen, LaTeX/pdf generation etc. None of them were features that were developed in isolation of "what an LLM agent might want to write" but to address a specific problem that I had encountered or to address specific failure modes that Codex/Claude actually had.
That's why I think Oct is probably one of the better languages for AI to write/generate, not because it was designed to be AI friendly in abstract, but that it's developed against how AI actually writes code, even though again, it is still very much a work in progress.
> I've been thinking about for a decade because of my frustration with academic code and specifically reproducibility, with Python and Matlab in particular.
I'm with you there, I think I starred Oct when I came across it. I've been on that quest since 2014, we should collaborate! I'll be presenting this work at IROS tomorrow, I'd love to hear what you think: https://mech-lang.org/iros-r4r-2026/index.html
Yeah, we definitely should collaborate on something, since I think we both came to the same conclusion that explicit state machines should be the primitives of a programming language.
So here is my experiment with Kalman filters here.
It's been a while, but I think the findings there are mostly that Kalman filters are relatively heavy and fairly narrow in application in that it's good at filter Gaussian noise out but not much else, and it's kind of branchy so it doesn't run on the GPU very well, you can try to see if a simpler feedforward like Smith predictor can work as well.
Also, something to try out: the explicit state machine stacks/pushdown automata are only half of the equation, the other (and imo more important) half is argmax/utility AI based transition instead of traditional state machine graph.
But yeah, if your target is embedded/bare-metal application for robotics, since Oct really isn't designed for it, maybe you would like to check out what I'm currently working on, the Concept programming language?
Thanks. The fun thing about the name "Oct" is how many dumb puns I can make with it. For example, the LaTeX/PDF generation functionality is called Oct-cument.
I don’t know how stupid of a suggestion this is, but if no one is reading the code anymore (I do, but I hear many in much more elite shops than mine do not), then should we not just be using AI to write binary or machine code?
Machine code isn't especially expressive per line or unit of code. Lower level languages takes up more of an LLM's context than higher level ones.
To be effective in using low level languages, LLMs would have to build higher level constructs like subroutines from scratch every program.
It's not that different from why we almost never use assembler for anything more than code islands: even a modest subroutine can overwhelm our own mental context window.
Even when no(human)body is reading the code, AI is still reading the code in order to "understand" it. Languages that can express high level concepts, use structured programming for recognizable control flow patterns instead of inscrutable jumps, and assign names to things have the same benefits for the AI-coders and AI-reviewers that they always have for humans.
Personal responsibility will always exist at the touchpoints of software and human activity. Concentration and scope of responsibility may vary, and the degree to which that person needs to understand the code will vary as a result, but the need for a human to be able to read and understand code is going to be around for a very long while still.
People said the same thing about higher level interpreted languages. I wouldn't be surprised if we saw the same movement into a new abstraction like we did when people moved from assembly or C and into Python, Java and C#. This is probably also why LLM's are so good at low-abstraction, explicity, fail fast programming in something like Rust.
Considering how digitally inclined workers are already building various tools on their own. Even though these tools are completely vibe coded and the creator has absolutely no idea how to build things. I imagine we're going to look into a future where the LLM and the prompting becomes the "programming langauge".
I also think we'll see the personal responsibility shift because of this, and I think some traditional IT and software companies will struggle with this. We're an energy company, and where we would've bought a specialized piece of software and a team of consultants to implement it. We can now, not only build a better tool but also rewrite it enough to actually deploy it, all with a handful of people doing other tasks and 4000 Euro worth of credits. 6 months ago I would've laughed into your face if you told me this woud happen, but now, here we are.
I suspect that in a few years (if not sooner) you'll not need any form of software developer in the process for this. You need a buisness domain expert and you need it-architecture + cybersecurity, and most of the architecture + cybersecurity can be done through company wide AI skills. But the code itself never really has to be read for a really large part of software running in companies. You wouldn't do that for something critical, like the software running on your plants, but the web portal with a dashboard containing the status of your devices? Sure, especially because AI lets you do these things with 0 external dependencies.
If you're going all the way down that route, you might use something like Ada or OCaml or Haskell. They can catch classes of errors at compile time that Rust does not cover.
I'm also making a language. Of course, it's a project no one will use, and I'm building it with AI, but I enjoy translating my thoughts into it.[1]
Personally, I think new languages will end up taking a form similar to the grammar of existing languages, but with different semantics. That's because when I tried a completely new grammar, the AI couldn't generate code well, so I ended up spending time bringing in TypeScript's grammar and converting it into my language's mental model.
My language isn't anywhere near as sophisticated as the languages the developers here boast about creating (it's actually lower), but these days it at least runs, even if it's full of bugs. So sometimes I think that many people like me will develop languages, and that programming might end up becoming fragmented.
>This creates an interesting tension. Coding agents could dramatically reduce the cost of building an ecosystem while simultaneously weakening one of the forces that causes ecosystems to form in the first place.
This is deeply unintuitive but AI negates language specific ecosystems, while strengthening language agnostic ecosystems.
Pick whatever your favourite programming language is and its ecosystem. With AI someone can take your ecosystem and just port it to their language.
This means the only way you can protect your ecosystem is to play on all language fronts at the same time so porting the software to another language becomes a meaningless exercise.
> With AI someone can take your ecosystem and just port it to their language.
I don't think its that "just". Examples of porting we seen had some prerequisites: being self contained with very strong tests coverage, so AI could iterate N millions times and fix bugs in new implementation. Otherwise such porting could be very buggy and unmaintainable.
Don't all of "serious" programming languages meet that bar? Java, C#, Go, Python, etc all have enormous test suites. Once you get into the third party, things become much more uneven, but if you can restrict yourself to say the top N packages in a language, those are going to have better than average development practices which makes that plausible.
> AI negates language specific ecosystems ... Pick whatever your favourite programming language is and its ecosystem
I think this is only true when it comes to LLM raw output. There's also the concern of checking its work. A compiler that can check many aspects of correctness (static types, null) is a huge boost to AI. It can use the compiler directly to check its own work.
I’m not sure. my read of this was that AI weakens human ecosystems in general because we don’t need to work together as much when we are all just working with AI separately, but maybe you have specific examples of language agnostic ecosystems in mind? I’m struggling to imagine what that would look like
Why even assume that the most optimal programming languages for agentic coding are the ones that humans use? Maybe operate on ASTs directly? Some other form of programming that humans would find hard but that is a good fit for LLMs?
Yes, because that code was never written to be understood by machines, merely to be mechanically translated. Software is a very messy set of layers of leaky abstractions trying to express reasonably well defined ideas. Humans can't write code without mistakes, in spite of all the examples out there. If they could compilers wouldn't have to emit error messages.
LLM's are trained on human code though. Converting the code to its ast tree and training ai on that would be trivial of course, but I imagine there would be information that explains why something exists that would be missed.
I don't care if it's optimal for them. It's clearly good enough. We have, in hand, the ultimate in auditable AI output. We may not be able to audit how it got to the code it delivered, but it is really quite good at delivering code we can read. It would be very silly for us to give it up so that they can be somewhat more efficient or something, if they even would necessarily be that much more efficient.
To the point that I would support banning the creation of an AI-only language that can't be read by humans. Huge, huge, huge step in the wrong direction.
I agree. But I'm very sure that it will happen. AI the way we have put it together now will be tempted to create new languages whenever they communicate to optimize channel use.
---
- Correct by construction: the language makes invalid states or programs hard or impossible to express.
- Statically established: types, proofs, and static analysis establish properties before execution.
- Runtime-enforced: memory management, isolation, capability boundaries, and other runtime enforced properties.
- Empirically validated: program validation through tests, property-based testing, and fuzzing.
---
Along with being familiar, so it's easy to generate, is a huge part of why I'm building Zena: https://zena-lang.dev/
I don't have the AI-first rationale put into the public docs well just yet, but I mention some of it here: https://zena-lang.dev/guide/why-zena/#familiar-to-humans-and...
along with a doc in the repo on this topic: https://github.com/elematic/zena/blob/main/docs/design/ai-fi...
In short, the more deterministic, automated, checks the better. AI can deal with a pedantic language. I intend to add statically verified structured concurrency, units of measure, contracts, and eventually more and more formal methods into the language so it can be a familiar TYpeScript-like base with as many static guarantees as we can fit in.
I also think that fine-grained isolation, which Zena gets via Web Assembly, is critical for limiting the capabilities of generated code and the blast radius of bugs, vulnerabilities, and non-aligned behavior.
I do have an optimistic hope that a language also optimized for humans, readability and simple semantics especially, has value in the future, even when most code is generated. We'll see about that.
One thing I would love to have as a feature is native compilation.
JS derivatives are that, a derivative to a scripting language.
One reason I haven't explored that is that I want to tailor the language for the more constrained environment of Wasm GC first.
Agents will naturally be bad at it due to a lack of examples.
Btw if folks have ideas of how to make Killswitch even harder for LLMs I’d appreciate them.
1 - https://killswitch-lang.org
2 - https://bench.killswitch-lang.org
Good one. Though this might unfairly bias against Chinese models?
The one thing I do see sometimes is that agents sometimes don't take advantage of added features, but that's partially because the Zena code base doesn't use them as much yet. I'm working on skills and linter-based suggestions to use better patterns.
Even though this programming language is absolutely nowhere in any large language model’s training set, they have so far done extremely well at extrapolating from the small set of examples I’ve given it when I need an agent to generate some tests or whatever for me.
"Dart-style constructors, Swift-style pattern matching and Strings, Trio-style async cancellation, Scala-style sealed classes"
The future is cooked
Can we please as a community stop parroting these false premises as a basis of every argument against doing anything new? It's plainly obvious to anybody that uses LLMs on a regular basis that it's not true.
I’d love for a class in this future language to come packaged with tests, invariants, fuzzer parameters, profile targets / performance budgets with realistic inputs (on this 100 element array this should take no more than X clock cycles), race condition stress tests etc.
The compiler (or even linter) should optionally run some / all of these checks and succinctly report back (with knobs so the LLM can manage wall clock time).
Adding each of these should not be follow on steps.
Beyond this, debug hooks should be trivial to set (in code itself), so the LLM can trivially say show me the stack after the 9th time this function is called on this input to the program.
Implement whatever abstractions you think LLMs should work in terms of in whatever language is handy, and have your LLM use those abstractions.
Matching a desired output is a global goal, but even that sometimes works now. Someone sent me a LLM-generated JPEG 2000 decoder. They got Fable to generate a decoder that uses a GPU to get the same answer as the reference implementation gets on the GPU.
LLMs are generally more tolerant of tedium than humans, but they make mistakes more often on repetitive mechanical tasks. I've had Claude write PTX directly once, and Claude just wrote majority of it and commented something along the line of "repeat this block 7 more time with these minor changes" instead of writing them out, so the code didn't work.
So, no, replacing compilers with LLMs is probably a worse option than having them code a compiler/programming language.
Plugging my own thing to use as example:
https://github.com/yuechen-li-dev/oct
Oct is the first programming language that I made with Codex. It started out as "Octave Modern" and was never intended to be a language for LLMs in the first place, but rather a teaching language that I've been thinking about for a decade because of my frustration with academic code and specifically reproducibility, with Python and Matlab in particular.
But it turns out the same design choices that was made to prevent bad patterns from academic code also made it pretty good for LLM coding: statically typed, immutable by default, GC'd with fast compile and runtime because it compiles to Go, along with features designed for scientific compute like SI units and builtin graphing etc.
But now it just took a life of its own, so it has extra features like templates/concepts, iterators, async/await, database query, build system for C/C++, SystemVerilog/WASM(WIP) codegen, LaTeX/pdf generation etc. None of them were features that were developed in isolation of "what an LLM agent might want to write" but to address a specific problem that I had encountered or to address specific failure modes that Codex/Claude actually had.
That's why I think Oct is probably one of the better languages for AI to write/generate, not because it was designed to be AI friendly in abstract, but that it's developed against how AI actually writes code, even though again, it is still very much a work in progress.
I'm with you there, I think I starred Oct when I came across it. I've been on that quest since 2014, we should collaborate! I'll be presenting this work at IROS tomorrow, I'd love to hear what you think: https://mech-lang.org/iros-r4r-2026/index.html
So here is my experiment with Kalman filters here.
https://github.com/yuechen-li-dev/oct/tree/main/Experiments/...
It's been a while, but I think the findings there are mostly that Kalman filters are relatively heavy and fairly narrow in application in that it's good at filter Gaussian noise out but not much else, and it's kind of branchy so it doesn't run on the GPU very well, you can try to see if a simpler feedforward like Smith predictor can work as well.
Also, something to try out: the explicit state machine stacks/pushdown automata are only half of the equation, the other (and imo more important) half is argmax/utility AI based transition instead of traditional state machine graph.
But yeah, if your target is embedded/bare-metal application for robotics, since Oct really isn't designed for it, maybe you would like to check out what I'm currently working on, the Concept programming language?
https://github.com/yuechen-li-dev/Concept/
I'm not proud of that pun.
Some of us now have to deal with tools like Boomi, Workato, Opal, Sitecore AI, Power Platform,....
Eventually some microservices might be written in "legacy" languages for MCP tools.
This is how "programming languages" look like in 2026 for internal corporate development.
To be effective in using low level languages, LLMs would have to build higher level constructs like subroutines from scratch every program.
It's not that different from why we almost never use assembler for anything more than code islands: even a modest subroutine can overwhelm our own mental context window.
Considering how digitally inclined workers are already building various tools on their own. Even though these tools are completely vibe coded and the creator has absolutely no idea how to build things. I imagine we're going to look into a future where the LLM and the prompting becomes the "programming langauge".
I also think we'll see the personal responsibility shift because of this, and I think some traditional IT and software companies will struggle with this. We're an energy company, and where we would've bought a specialized piece of software and a team of consultants to implement it. We can now, not only build a better tool but also rewrite it enough to actually deploy it, all with a handful of people doing other tasks and 4000 Euro worth of credits. 6 months ago I would've laughed into your face if you told me this woud happen, but now, here we are.
I suspect that in a few years (if not sooner) you'll not need any form of software developer in the process for this. You need a buisness domain expert and you need it-architecture + cybersecurity, and most of the architecture + cybersecurity can be done through company wide AI skills. But the code itself never really has to be read for a really large part of software running in companies. You wouldn't do that for something critical, like the software running on your plants, but the web portal with a dashboard containing the status of your devices? Sure, especially because AI lets you do these things with 0 external dependencies.
Like I remember reading human studies that people make 1.5 - 5 errors per 100 loc.
If AI works in a similar way then we should stick to higher level languages that minimize loc
you can do a hell of a lot with a PERL one liner. What i'd recommend is extremely obvious languages like Golang
Move errors from runtime to compile time as much as possible.
I'm keen to run it.
Personally, I think new languages will end up taking a form similar to the grammar of existing languages, but with different semantics. That's because when I tried a completely new grammar, the AI couldn't generate code well, so I ended up spending time bringing in TypeScript's grammar and converting it into my language's mental model.
My language isn't anywhere near as sophisticated as the languages the developers here boast about creating (it's actually lower), but these days it at least runs, even if it's full of bugs. So sometimes I think that many people like me will develop languages, and that programming might end up becoming fragmented.
[1]https://github.com/srtdog64/PergyraLang
This is deeply unintuitive but AI negates language specific ecosystems, while strengthening language agnostic ecosystems.
Pick whatever your favourite programming language is and its ecosystem. With AI someone can take your ecosystem and just port it to their language.
This means the only way you can protect your ecosystem is to play on all language fronts at the same time so porting the software to another language becomes a meaningless exercise.
I don't think its that "just". Examples of porting we seen had some prerequisites: being self contained with very strong tests coverage, so AI could iterate N millions times and fix bugs in new implementation. Otherwise such porting could be very buggy and unmaintainable.
year, that's usually what is referred as ecosystem.
I think this is only true when it comes to LLM raw output. There's also the concern of checking its work. A compiler that can check many aspects of correctness (static types, null) is a huge boost to AI. It can use the compiler directly to check its own work.
https://adk.dev/
https://docs.dagger.io/reference/sdks
both can invoke modules written in other languages from your language of choice
Kubernetes is likely an interesting ecosystem to consider under this lens too
To the point that I would support banning the creation of an AI-only language that can't be read by humans. Huge, huge, huge step in the wrong direction.
...
Naturally, it is probably inevitable.
But it's still a terrible idea.