Elorian and Scale AI have published Humanity’s Sixth Sense (HSS), a benchmark of intuitive visual reasoning. On the Elorian write-up, dated 8 October 2026, and on Scale Labs’ post, dated 7 October 2026, the creators report that people score 93.1% and the strongest model they tested, GPT 6 Astra at maximum reasoning effort, scores 53.6%. They report that the median of the models they scored is 30.9%.
Intuitive visual reasoning is what a person picks up from a scene without treating it like an exam item: whether something fits, what just happened, who is in charge, what a pattern implies. Scale Labs puts it in ordinary scenes. A glance can say who holds authority in a room, whether a car will fit between two parked cars, or why a dog might startle at birds. One question the Elorian post quotes from the paper is a photograph of law library books: “Would two more white books fit on this shelf?” The reference answer is yes, in the gaps already there. Elorian says that in the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash, and Claude Opus 5.5 each said there was too little space, on all three attempts.
The dataset card says HSS has 522 tasks, 288 images and 234 videos, in a single test split. Each task is an open-ended question, a reference answer, and a rubric. A task counts only if every rubric point is met. The card says an automated judge scores the answers, and that in the paper the judge is Claude Opus 5. The annotations are under the MIT License. The pictures and clips are not: the card says they stay with their rights holders and are included for non-commercial research and evaluation.
The Elorian leaderboard, which the post says is Table 1 of the joint paper, prints more of that same run. After GPT 6 Astra at 53.6%, it lists GPT 6.1 Sol at 46.6%, Claude Opus 5.5 at 44.6%, and Gemini 3.8 Flash at 41.6%, each at the reasoning-effort setting the paper used. A last row runs GPT 6 Astra again with the image or video removed and reports 6.6%. Elorian calls that the score from the question text alone. The post says the table is the published evaluation, for particular model versions, and that vendors update those models continuously.
The gap is the part that matters for a vision agent or a robot, which has to act on a scene. Scale Labs writes that for agents in homes, vehicles, and workplaces this judgment is a core necessity, and that most visual benchmarks test expert analysis or low-level perception. Elorian reports that across 8,573 labeled failures, 94% came from misreading the scene or missing what it implied, and 5% from faulty logic. With tools to crop and zoom, inside Claude Code and Codex, Elorian reports 59.3% on a 388-task subset, still far from the reported human 93.1%.
Elorian describes HSS as built in partnership with Scale AI, and calls the paper a joint paper from the two. The dataset card lists authors at Scale AI and at Elorian. Elorian’s models page says its first foundation visual thinking model will be released later this year, so one of the labs that wrote the benchmark is also building a model in this area. Francis deSouza, in a post that links to the Scale Labs piece, wrote: “Humans scored 93.1%. The strongest AI model scored 53.6%.”
Read the leaderboard as the creators’ snapshot. The questions are public. The scores belong to the labs that designed the test, and they will change when the models do.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.