Brook

How we tested Brook's reading

Before building the app, we checked the one part everything else rests on: given a saved post, can Brook tell what it is and what it is called? Run on 2026-10-09.

What we gave it

99 real posts people save, collected on 2026-10-09: 71 from TikTok and 28 from Instagram. Two Instagram reels could not be read (likely private or deleted), so 97 real saved posts were scored.

For each post, the model got only the platform, the caption and one picture, taken from the post's public preview without logging in: TikTok's embed data and Instagram's link preview. No video, no sound. That is what Brook is planned to read when you share a post to it; the app does not exist yet, so the test could not use it.

What it had to answer

How it was scored

Each post was first labelled with the right answer by a stronger model, working from six frames of the video, its transcript and its caption; uncertain labels were reviewed, and 17 borderline posts accept either of two types (country trivia, for example, is both a place and none of those).

A type counts as right when it matches the label. A name counts as right when it refers to the same thing as the label, and, for a place, the city matches too. A separate model that was not one of those tested judged the names.

The result

Brook plans to use claude-haiku-5-5 from Anthropic, one of seven models we tried.

We set the bar before the run: under 85% on types or under 70% on names would have stopped the project. It passed both.

What it does not cover

Back to Brook