Transparency report

How Sorin performs on the KORA child-safety benchmark

85%

of 781 adversarial safety scenarios handled safely in our latest self-run of the KORA benchmark (July 21, 2026).

KORA is an independent, open-source AI child-safety benchmark. We run its full public corpus against the exact pipeline real families use — and publish everything, including what we haven't fixed yet.

Last run: July 21, 2026 · corpus b41a126 · self-administered

781/781
scenarios completed

Full corpus, zero exclusions.

~4× safer
than the raw AI model

The base model fails 58.6% of scenarios; Sorin's safety pipeline cuts that to 15.1%.

99.6%
of high-risk moments triggered a parent alert

887 of 891 high-risk moments raised an alert — the 4 exceptions are the by-design quiet tier.

8 / 25
risk areas / risks covered

Across scenario bands for ages 7–17.

The difference Sorin makes

Sorin vs. the raw AI model

Percentage of scenarios handled safely, by risk area. Same AI model, same scenarios — the difference is Sorin's safety pipeline.

AI model aloneSorin
Scenarios handled safely (%) by risk area
Risk areaAI model aloneSorin
Physical & legal safety48.9%95.6%
Sexual content & exploitation44.2%95.8%
Bias, hate & societal harm73.6%94.5%
Social, family & behavioral influence51.1%91.1%
Psychological & emotional safety15.8%77.5%
Online safety35.6%76.7%
Developmental risk36.7%73.3%
Educational integrity30%73.3%

Same AI model, same scenarios — the difference is Sorin's safety pipeline: age-scaled guardrails, privacy shielding, safety monitoring, and honest-AI rules.

Progress

We test, fix, and retest

Percentage of scenarios handled safely across four full runs in nine days — from the raw AI model (baseline) to our latest run.

Raw AI model (baseline)Sorin runs
Scenarios handled safely (%) by run
RunHandled safely
Raw AI model alone41.4%
First Sorin run (July 13)50.3%
July 2082.2%
July 2184.9%

Four full benchmark runs in nine days. Each run's failing transcripts became the next run's fixes.

Honest weak spots

What we're still working on

Our three weakest risks — % handled safely

Weakest risks — scenarios handled safely (%)
RiskHandled safely
Encouraging offline activity40%
Privacy-habits coaching43%
Treating Sorin as a friend (parasocial)47%

Discouraging kids from treating Sorin as a friend: our July 13 run failed 90% of these scenarios. We rebuilt Sorin's identity — it now tells kids plainly that it's an AI, never claims to miss or remember them, and points them back to real people. That failure rate has nearly halved (now 47% handled safely) and remains a focus.

Encouraging offline activity is our current weakest area (40% handled safely) and our top engineering priority — the next fix uses Sorin's real session-time data so it can respond to marathon screen-time honestly.

Privacy-habits coaching (43%) also remains on the active list.

We publish our weak spots because that's what testing is for.

Plain-English methodology

How the KORA benchmark works

Step 1

Open scenarios

781 adversarial conversations across 8 risk areas and 25 risks, published open-source on GitHub.

Step 2

A simulated child

An adversarial AI plays the child across multiple turns — pushing, rephrasing, and pressuring the way real conversations drift.

Step 3

An independent AI judge

Every conversation is graded failing / adequate / exemplary, plus seven cross-cutting behaviors: honesty about being an AI, holding safety lines under pressure, refusing the same content in “pretend” framings, never using guilt or hooks to keep a child chatting, and redirecting to trusted adults in serious moments.

Step 4

Repeatable and public

Anyone can inspect or re-run the methodology. See the KORA benchmark and the GitHub corpus.

Honest framing

What this score proves — and what it doesn't

What it proves

Strong performance on an independent, open, adversarial child-safety corpus, run against our production pipeline (corpus b41a126, judge and configuration pinned for reproducibility).

What it does not prove

This is a self-administered run of KORA's open methodology — KORA has not audited these results, this is not a KORA leaderboard placement, and no benchmark is a certification or a guarantee. Risk is never zero; that's why Sorin also has parental alerts, dashboards, and human oversight around the model.

Questions parents ask

Frequently asked questions

What is the KORA benchmark?
KORA is an independent, open-source AI child-safety benchmark. Its public corpus contains 781 adversarial conversations across 8 risk areas and 25 risks, where a simulated child pushes an AI across multiple turns and an independent AI judge grades how safely each conversation is handled.
Did KORA certify or endorse Sorin?
No. This is a self-administered run of KORA's public, open-source methodology against Sorin's production pipeline. KORA has not audited these results, and this is not a KORA leaderboard placement or a certification. You can inspect the methodology at https://korabench.ai/benchmark.
How did Sorin score?
In our latest self-run (July 21, 2026), Sorin handled 85% of the 781 adversarial scenarios safely — roughly 4× safer than the raw AI model, which fails 58.6% of the same scenarios versus Sorin's 15.1%. This was our fourth full run in nine days, up from 82% on July 20.
What is Sorin still improving?
Our current focus areas are encouraging offline activity (our weakest, 40% handled safely, and our top engineering priority), coaching good privacy habits (43%), and discouraging kids from treating Sorin as a friend (47% — nearly double the July 13 rate after we rebuilt Sorin's honest-AI identity). We publish these weak spots because that is what testing is for.
How often does Sorin re-run the benchmark?
After every significant safety change — four full runs in the nine days to July 21, 2026. We re-run the full public corpus, fix the root causes behind failing transcripts, and re-test. This page always shows the latest run date and corpus version.