AI / Hikaru Series / Model
Hikaru-1.
Our first model, and the testbed for Explanatory Reinforcement Learning. Still in development, at a reduced pace, while Deriva-1 gets most of our attention.
Overview
Hikaru-1 is a finetune of Ministral-3 (8B),
trained with our Explanatory Reinforcement Learning (ERL) method.
It exists to answer a single question: does a model trained with
explanatory reinforcement learning behave measurably differently
than one trained with a binary reward? Until that question has a
measured answer, larger ambitions would be built on assumption.
Hikaru-1 removes the assumption.
Hikaru-1 entered development on June 16, 2026. After we reviewed how many projects we had running at once, we set a clear order of priority across the company: Kova first, then Deriva-1, then Hikaru-1, then PICO-1. Hikaru-1 continues under that plan, just at a reduced pace rather than paused. See Evaluating scope for the full reasoning.
Capabilities
What Hikaru-1 can do.
Hikaru-1 is a small, focused validation model. Capabilities not listed here have not been finalized, and we would rather leave them out than guess.
Benchmarks
Not published yet.
Benchmark runs are not finished. We would rather leave this table empty than publish numbers we can't stand behind, so every score below is marked N/A until that changes.
| Benchmark | Hikaru-1 | Ministral-3 (8B) (base) | Comparison |
|---|---|---|---|
| General reasoning | N/A | N/A | N/A |
| Instruction following | N/A | N/A | N/A |
| ERL vs. binary reward (internal) | N/A | N/A | N/A |
| Safety / refusal balance | N/A | N/A | N/A |
Sential Innovations is an independent community and is not affiliated with, sponsored by, or endorsed by the makers of the base model. Any release will follow the base model's original licence, with attribution, and will be provided free and as-is.