Good enough for daily use, not ready for release
Today we're marking a real milestone: Kova has moved into
internal beta. That means it's good enough for us to use
day-to-day, not just to test in isolated demos. It does not
mean Kova is release-ready. It is still very rough around the
edges, and there is a long list of polish, stability, and safety
work between here and a public launch. But reaching daily-use
quality internally is proof the core approach works, and it has
opened the door to testing capabilities we could not
responsibly test before.
Why we're testing under tighter restrictions than we'll ship with
Kova is powerful, and power without care around it is a
liability. It can act with real consequences: spend money,
control your home, browse the web, write and run code. Because
of that, we are deliberately testing it under more restrictive
settings internally than we plan to ship publicly. That sounds
backwards, but it isn't: it lets us find where the restrictions
actually matter before we decide how to relax them, so the
version people eventually get is capable and safe by default,
not safe because it's crippled.
The models behind this round of testing
Internal beta testing has run across a wide mix of models, not
just our own, so we can see how Kova performs regardless of
what's doing the reasoning underneath. So far, that list
includes:
Deriva-0.2 (non-release)
Hikaru-0.1 (non-release)
MiniMax-M3
Claude Opus 5
Fable 5
GPT-5.6 SOL
Ornith-1.5:9b (via the developer integration)
Grok-4.6
Llama 3.1
- Several additional internal models we can't disclose yet
We went into this round expecting Kova to be usable. What
surprised us was how far it could be pushed, even under
restriction, across such a wide range of models.
What comes next
Internal beta continues, and later updates will cover what
changes as we work through stability, safety, and the rest of
the road to a public release. Stay tuned.