neurovidz · neurovideo
trust · 2026-07-14

Can you trust an engagement score? How we validate NeuroVidz.

It’s the right question, and it deserves a real answer — not a badge that says “99% accurate.” For creative work there is no single ground truth to be 99% accurate against: the same clip can land with one audience and die with another. So instead of publishing a number we couldn’t defend, we’ll show you how the system is built to be hard to fool, hard to break, and easy to falsify yourself.

1. The foundations are published science, not vibes

The read maps perceptual signals onto a seven-network model of brain organizationthat comes from peer-reviewed neuroscience — the widely used cortical parcellation established by Yeo, Krienen and colleagues (2011), the standard reference for how attention, salience, and default-mode systems divide the cortex. The engagement composite leans on findings that have survived replication: sustained attention opposes mind-wandering systems, emotionally arousing content is shared and remembered more (Berger & Milkman’s well-known virality research), and narrative immersion recruits default-mode regions during naturalistic viewing. Our methods page walks through the full pipeline.

2. Every scoring change must survive real clips

Aggregate metrics lie. A change can look better “on average” and read worse on the clips that matter. So every change to the scoring or emotion pipeline runs against a regression suite of real material — long-form music, spoken audio, silence, monotone speech, fast-cut video — and has to preserve or improve the read on each, not just the mean. We keep what we call a falsification log: changes we designed, tested, and rejectedbecause real clips got worse. More than one “improvement” that looked great on paper died there. The log is why the product moves slower than a demo and breaks less than one.

3. The boring engineering that keeps reads honest

Two implementations of the emotion decode exist — one server-side, one in your browser — and they are checked to produce bit-identical output, so what you see is exactly what was computed. Timelines carry hard guards: if any component produces timestamps that don’t cover the clip or arrive in the wrong units, the system rejects that read and falls back rather than stretching data to fit. (That guard exists because our own verification once caught a timeline arriving in minutes instead of seconds — it never reached users, and now it never can.)

4. When it isn’t sure, it says so — and it costs you nothing

Weak signal produces “no clear read,” not a confident-sounding guess. Every emotion segment ships with the evidence phrase behind it. The engagement score shows its components — measured signal, content judgment, heard emotional arc — instead of hiding behind one mystery number. And as of this week: if an analysis fully abstains, its credits are refunded automatically. A system that charged you for shrugging wouldn’t deserve trust.

5. Built in-house — with one honest exception

The scoring engine — signal extraction, the brain-network mapping, engagement and emotion scoring — is ours, built and calibrated in-house on the published literature above. The one thing we don’t build ourselves: the plain-language layer. Summaries and recommendations are written with a commercially licensed language model under its commercial terms, fed by our measurements. It phrases; it does not score.

6. Falsify it yourself in ten minutes

Don’t take our word for any of this. Pick two clips you already know the truth about — your best performer and one that flopped — and run both. A trustworthy read should separate them, locate the moments you already know mattered, and admit uncertainty where your material is genuinely ambiguous. If it doesn’t, tell us — the falsification log has room. You can also inspect a full sample result without creating an account.

try it on your own clip

Every new account starts free — the founding 50 get 40 credits (+ a founding badge), then 20 for everyone after. No card required.

Start free →