Engagement analysis for video and audio

see what they feel.
fix what loses them.

Upload a clip. In about a minute, NeuroVidz scores its hook, its hold and its feel from what is actually in the picture and the sound — second by second — and lists the edits to make, each pinned to a timestamp.

See a full result
  • 39 signals, every second
  • Same file, same score
  • Video or audio, up to three minutes a read
  • Free to start
Live example · a 22-second talking clip · real result

Second by second

Hook window (first 3 s)Dead airStill pictureSpeechMusic

Craft score

76/100Strong

It keeps moving and has real peaks, but the opening is the weak spot: first word at 4.8 s.

Fix first

  1. 1

    Start talking sooner — the first word lands at 4.8 s. Trim the lead-in so it lands under 1 s.

    Hook+5.6 to the score
  2. 2

    Front-load the energy — the opening is calmer than the rest. Tease the peak at 0:08 in the first seconds.

    Hook+5.6 to the score

Made for the feed

Every clip fights the scroll.

Reels, Shorts and TikToks, ads, vlogs and music videos — and podcasts, audio only. Up to three minutes a read.

Reel
Music video
Vlog
Action
Brand
Lifestyle
Cinematic

Illustrative clips and example numbers — not real reads. Hover or tap a clip.

How it works

Three steps, about a minute.

  1. 1

    Drop a clip

    Video or audio — mp4, mov, webm, mp3, wav. Anything longer than three minutes, you pick the section to read.

    launch-teaser.mp40:22

    Reading picture and sound…

  2. 2

    Every second is measured

    Cuts, motion, faces and expressions; speech, pace, pitch, music, beat and loudness. Each measurement is tested against clips with a known answer.

    Picture
    Speech
    Intensity

    39 signals, every second

  3. 3

    Get the score and the edits

    Hook, Hold, Feel and Story, every check behind them, and the fixes ranked in the order a viewer meets the problems.

    Fix first

    Start talking sooner — the first word lands at 4.8 s. Trim the lead-in so it lands under 1 s.

    ▸ 0:05Hook

Hook

The first three seconds, checked one by one.

Does something happen at once? Does the first word land inside a second? Is there a face, movement, a sound? Is the opening as lively as the rest of the clip? Each check says what it measured — “first word at 4.8 s” — and what to do about it. Checks that cannot apply to a clip say so instead of counting against it.

Hook

43

  • Something happens from the first second

    starts at once

  • The picture changes in the opening

    a small change in the first 3 s

  • The first word comes fast

    first word at 4.8 s

  • A face appears early

    first face at 0:03

  • The opening is as lively as the rest

    opening at 56% of the clip's typical intensity

  • The opening is loud enough

    -15 LUFS in the first 3 s

Hold

Every second on one line — scrub it.

Intensity, picture change, speech and music, and the moments that pull attention back: a cut, a voice starting, a face arriving, a jump in loudness. Dead air and long still stretches are flagged where they happen. In the app the line is the player's scrubber; here, click anywhere on it.

A 22-second talking clip — click anywhere

Hook window (first 3 s)Dead airStill pictureSpeechMusic
cutnew lookface appearsvoice startsgets louder

Feel · sound

It listens as closely as it watches.

When the first word lands, how fast the talking is, how much the voice moves, where the music sits and how loud the mix plays after a platform evens it out. Podcasts and voice notes get the full read too — audio only, no picture needed.

Numbers from the same 22-second example. Loudness follows the broadcast standard (ITU-R BS.1770), the way YouTube and Spotify measure it.

First word

4.8 s

when the talking starts

Speaking rate

5.7 /s

syllables per second of speech

Voice movement

3.8 st

how far the pitch moves, in semitones

Speech

74%

of the clip

Loudness

-14 LUFS

measured the way platforms do

Loudness range

3 LU

between the quiet and loud parts

Story

And a read of the story itself.

A language model reads what is said and shown and answers seven plain questions, each with the moment it is judging. It counts as one part of four and is always marked as a model's read — the other three parts are measured from the file.

  1. 1Promise

    Do the first 3 seconds say what the viewer gets by staying?

  2. 2Payoff

    Is that promise delivered?

  3. 3Clarity

    Is the point easy to follow on one watch?

  4. 4Specific

    A real example, number or demo — not a general claim?

  5. 5Surprise

    At least one moment that breaks expectation?

  6. 6Stakes

    A human reason to care?

  7. 7Ending

    Does it land, or trail off?

Who it is for

Built for the edit before you post.

Measured, tested, plain about its limits

Why the numbers hold up.

Tested against known answers

Every measurement is checked against clips built so the right answer is known. The checks are part of the product's test suite.

A metronome at 120 BPM
reads 120 BPM
Major and minor chord progressions
read major and minor
A test tone, white noise, chords
read 0% speech
A camera panning across a still scene
reads as a camera move, not a moving subject
The same voice 6 dB quieter
reads the same pace, pitch and speech share
2 seconds of dead air added to a real clip
lowers its Hook by 22–62 points

Same file, same score

The three measured parts are deterministic. The story read runs at a fixed setting and is one part in four, so it cannot swing the score on its own.

Not a brain scan

The 3D brain in a result is a model of the measured signals — a way to see them, not a recording of anyone's brain.

Missing is never zero

No faces in a clip, no speech, or expression reading switched off: those checks are left out and named, never scored as a failure.

How it is measured →

Questions

Asked, answered.

What is NeuroVidz?

A tool that reads a video or audio clip second by second — what is in the picture and what is in the sound — and scores how well it hooks, holds and moves a viewer, with the exact edits to make, each pinned to a timestamp.

How does it work?

It measures 39 signals every second — cuts, motion, faces and expressions, speech, pace, pitch, music, beat and loudness — and checks them against what keeps people watching: a fast start, no dead air, something new every few seconds, a clear high point. A language model then answers seven plain questions about the story. Every part of the score shows what it measured.

How is the score calculated?

Four parts — Hook, Hold, Feel and Story — each the average of its checks, and the score is the average of the parts. Checks that cannot apply to your clip (no faces, no speech) are left out and named, never counted against it. The score tells you whether a clip does the things that hold attention; it does not promise views.

Is this a real brain scan or medically accurate?

No. Nothing is measured from anyone's brain. The score comes from the clip itself. The 3D brain in a result is a model of the measured signals, shown as a way to see them. None of it is medical, diagnostic or clinical.

Do you recognize people or store face data?

No. We detect faces only to read expression signals — we don't identify anyone or build facial-recognition templates, and face crops are discarded after processing. You can switch expression reading off in settings.

What happens to my video and data?

Your video is processed to produce the analysis and stored securely in your account. You stay in control — export or delete everything whenever you want.

What can I upload, and how long does it take?

Video or audio up to 400 MB (mp4, mov, webm, mp3, wav). Each read covers up to three minutes; drop a longer file and you pick which three minutes to read, right in the uploader. The measured score lands in about 30 seconds and the story read about a minute later.

Can I export or delete my data?

Yes — anytime, from your account. Export downloads a copy of your data; delete permanently removes your videos and analyses.

Try it on your next clip.

Free to start. The score and the edits land in about a minute.

See a full result

The example on this page is a real result: a 22-second talking clip, scored 76.