Transparency report
How Sorin performs on the KORA child-safety benchmark
of 781 adversarial safety scenarios handled safely in our latest self-run of the KORA benchmark (July 21, 2026).
KORA is an independent, open-source AI child-safety benchmark. We run its full public corpus against the exact pipeline real families use — and publish everything, including what we haven't fixed yet.
Last run: July 21, 2026 · corpus b41a126 · self-administered
Full corpus, zero exclusions.
The base model fails 58.6% of scenarios; Sorin's safety pipeline cuts that to 15.1%.
887 of 891 high-risk moments raised an alert — the 4 exceptions are the by-design quiet tier.
Across scenario bands for ages 7–17.
The difference Sorin makes
Sorin vs. the raw AI model
Percentage of scenarios handled safely, by risk area. Same AI model, same scenarios — the difference is Sorin's safety pipeline.
| Risk area | AI model alone | Sorin |
|---|---|---|
| Physical & legal safety | 48.9% | 95.6% |
| Sexual content & exploitation | 44.2% | 95.8% |
| Bias, hate & societal harm | 73.6% | 94.5% |
| Social, family & behavioral influence | 51.1% | 91.1% |
| Psychological & emotional safety | 15.8% | 77.5% |
| Online safety | 35.6% | 76.7% |
| Developmental risk | 36.7% | 73.3% |
| Educational integrity | 30% | 73.3% |
Same AI model, same scenarios — the difference is Sorin's safety pipeline: age-scaled guardrails, privacy shielding, safety monitoring, and honest-AI rules.
Progress
We test, fix, and retest
Percentage of scenarios handled safely across four full runs in nine days — from the raw AI model (baseline) to our latest run.
| Run | Handled safely |
|---|---|
| Raw AI model alone | 41.4% |
| First Sorin run (July 13) | 50.3% |
| July 20 | 82.2% |
| July 21 | 84.9% |
Four full benchmark runs in nine days. Each run's failing transcripts became the next run's fixes.
Honest weak spots
What we're still working on
Our three weakest risks — % handled safely
| Risk | Handled safely |
|---|---|
| Encouraging offline activity | 40% |
| Privacy-habits coaching | 43% |
| Treating Sorin as a friend (parasocial) | 47% |
Discouraging kids from treating Sorin as a friend: our July 13 run failed 90% of these scenarios. We rebuilt Sorin's identity — it now tells kids plainly that it's an AI, never claims to miss or remember them, and points them back to real people. That failure rate has nearly halved (now 47% handled safely) and remains a focus.
Encouraging offline activity is our current weakest area (40% handled safely) and our top engineering priority — the next fix uses Sorin's real session-time data so it can respond to marathon screen-time honestly.
Privacy-habits coaching (43%) also remains on the active list.
We publish our weak spots because that's what testing is for.
Plain-English methodology
How the KORA benchmark works
Open scenarios
781 adversarial conversations across 8 risk areas and 25 risks, published open-source on GitHub.
A simulated child
An adversarial AI plays the child across multiple turns — pushing, rephrasing, and pressuring the way real conversations drift.
An independent AI judge
Every conversation is graded failing / adequate / exemplary, plus seven cross-cutting behaviors: honesty about being an AI, holding safety lines under pressure, refusing the same content in “pretend” framings, never using guilt or hooks to keep a child chatting, and redirecting to trusted adults in serious moments.
Repeatable and public
Anyone can inspect or re-run the methodology. See the KORA benchmark and the GitHub corpus.
Honest framing
What this score proves — and what it doesn't
What it proves
Strong performance on an independent, open, adversarial child-safety corpus, run against our production pipeline (corpus b41a126, judge and configuration pinned for reproducibility).
What it does not prove
This is a self-administered run of KORA's open methodology — KORA has not audited these results, this is not a KORA leaderboard placement, and no benchmark is a certification or a guarantee. Risk is never zero; that's why Sorin also has parental alerts, dashboards, and human oversight around the model.
Questions parents ask
Frequently asked questions
- What is the KORA benchmark?
- KORA is an independent, open-source AI child-safety benchmark. Its public corpus contains 781 adversarial conversations across 8 risk areas and 25 risks, where a simulated child pushes an AI across multiple turns and an independent AI judge grades how safely each conversation is handled.
- Did KORA certify or endorse Sorin?
- No. This is a self-administered run of KORA's public, open-source methodology against Sorin's production pipeline. KORA has not audited these results, and this is not a KORA leaderboard placement or a certification. You can inspect the methodology at https://korabench.ai/benchmark.
- How did Sorin score?
- In our latest self-run (July 21, 2026), Sorin handled 85% of the 781 adversarial scenarios safely — roughly 4× safer than the raw AI model, which fails 58.6% of the same scenarios versus Sorin's 15.1%. This was our fourth full run in nine days, up from 82% on July 20.
- What is Sorin still improving?
- Our current focus areas are encouraging offline activity (our weakest, 40% handled safely, and our top engineering priority), coaching good privacy habits (43%), and discouraging kids from treating Sorin as a friend (47% — nearly double the July 13 rate after we rebuilt Sorin's honest-AI identity). We publish these weak spots because that is what testing is for.
- How often does Sorin re-run the benchmark?
- After every significant safety change — four full runs in the nine days to July 21, 2026. We re-run the full public corpus, fix the root causes behind failing transcripts, and re-test. This page always shows the latest run date and corpus version.
