CogniRoute Beats Top Proprietary Model by 15 Points on Social Video QA, Releases 118K Benchmark
Summary
Most multimodal benchmarks test factual recall; social video QA—where the answer hinges on a misread gesture, vocal sarcasm, or audio-visual mismatch—is substantially harder and closer to real human-AI interaction needs; a 15-27pp gain over proprietary models while releasing a 118K training benchmark is a meaningful step.
Originally reported by paper
Read the original article →Original headline: CogniRoute Beats Top Proprietary Model by 15 Points on Social Video QA, Releases 118K Benchmark