Hikaru-1
First model in the series
A finetune of IBM Granite 4.1 (8B), trained with Explanatory Reinforcement Learning. The first concrete test of whether ERL produces a measurably different model.
Hikaru-1 exists to answer a single question: does a model trained with explanatory reinforcement learning behave measurably differently than one trained with a binary reward? Until that question has a measured answer, larger ambitions would be built on assumption. Hikaru-1 removes the assumption.
- BaseIBM Granite 4.1 (8B), finetuned
- MethodExplanatory Reinforcement Learning (ERL)
- StatusIn development since June 16, 2026
- PurposeValidate ERL before scaling to larger models
Read the development announcement →
Sential Innovations is an independent community and is not affiliated with, sponsored by, or endorsed by the makers of the base model. Any release will follow the base model's original licence, with attribution, and will be provided free and as-is.