Scale Labs and Elorian release Humanity’s Sixth Sense benchmark
Scale Labs and Elorian released a 522-task benchmark that tests whether multimodal models can infer temporal, causal, spatial and social information implied by images and video.
Scale Labs and Elorian released Humanity’s Sixth Sense, an open-ended benchmark that tests whether multimodal models can infer information implied by images and video rather than only what is directly visible.
The release gives researchers 522 tasks built from 288 images and 234 video clips totaling 17.6 hours. Scale Labs also published the dataset and an evaluation harness with model, evaluation-prompt and grading-prompt resources for testing additional models.
Humanity’s Sixth Sense, or HSS, splits its questions across four domains and 11 subdomains. Temporal and causal tasks ask models to reconstruct an unshown past event, identify the physical cause of an ongoing event or anticipate what is likely to happen next. Physical and spatial tasks cover hidden properties, possible actions, unseen viewpoints and whether a target is reachable under spatial constraints. Social questions target beliefs, knowledge gaps, emotions, intentions, roles, norms and power dynamics. The fourth domain tests abstract patterns and consequences implied by a scene.
The authors say the tasks address gaps in conventional multimodal evaluations, which they characterize as emphasizing academic problem-solving, visible-object perception, narrow visual skills, synthetic settings or multiple-choice answers. According to the accompanying paper, HSS pairs each image or video with a human-written question, a reference answer and an atomic grading rubric. Models respond in free form, and an LLM judge counts a task as correct only if every rubric criterion is satisfied.
Scale Labs said 3,466 authored tasks passed through three independent review rounds, with 522 ultimately accepted — a 15.1% end-to-end acceptance rate. The company said accepted tasks had to require inference beyond visible details, avoid specialized expertise and receive unanimous reviewer agreement. For related benchmark context, see Scale Labs’ SWE-Bench Pro V2 benchmark update.
The benchmark makers reported a 93.1% human pass rate and a 53.6% pass@1 score for the highest-scoring tested system, GPT-6-astra at maximum reasoning effort. They put the median score across evaluated models at 30.9%. Those figures are issuer-reported results for specific model versions, prompts, settings and an LLM-based judge; the supplied research found no independent replication or audit.
The authors said they evaluated 25 models from eight vendors alongside a baseline drawn from 20 human annotators, using three attempts per task and averaging pass@1 across those attempts. They reported that 23 of the 25 models scored lower on video than on images, by 7.3 percentage points on average. Social understanding was the lowest-scoring domain for 21 models, averaging 24.4%, versus 34.1% across the other three domains.
For a blind control, the authors removed the visual input from GPT-6-astra and reported that its score dropped from 53.6% to 6.6%. They presented the result as evidence that the questions alone did not explain the scores. In a separate analysis of 8,573 labeled first-attempt failures, the team attributed 53% to perception, 41% to missing implied information and 5% to faulty reasoning.
The dataset card records an October 8 update that added evidence-window fields for all 234 video tasks and revised one prompt after an internal audit. The MIT License applies to annotations, but not to the collected images and videos. Those remain the property of their original rights holders and are included for non-commercial research and evaluation.
More news

Liquid AI releases d1 models for single-pass edge decisions

Google releases EmbeddingGemma 2 for on-device multimodal search

Mistral launches Large 4 in public preview
