HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllPapersBenchmarks

Home / Research / Benchmarks

Research / Benchmarks 6 items

All Benchmarks items under Research on AIFIRST News, newest first. 6 items.

Artificial Analysis "Intelligence": A meaningless benchmark

A critique argues that the Artificial Analysis "Intelligence" benchmark is fundamentally flawed, using Qwen 3.8 27B as a case study where benchmark scores diverge sharply from coding task performance. The post includes side-by-side comparisons showing where high benchmark rankings do not reflect real-world output quality.

/23 Aug 2026/reddit

ai-crawler-registry: machine-readable registry of AI crawlers with measured block rates

A machine-readable registry of more than 50 AI web crawlers, published as ai-crawlers.json by Crawl Census. Each entry carries robots.txt tokens, user-agent strings, crawler purpose, and compliance status, with block-rate data measured across more than 4,194 domains. The dataset separates training crawlers from answer-engine and retrieval crawlers run by the same operator, including OpenAI, Anthropic, and Google, and records which agents state that they disregard robots.txt. Licensed CC BY 4.0 with a live API.

/22 Aug 2026/GitHub

Seedance 2 has already been dethroned

Alibaba's "Happy Horse" model (open source, expected) surpassed Seedance 2 on image-to-video leaderboard. Low-engagement post, possible promotional content (no audio support noted by commenter, community skeptical of ...

/16 Apr 2026/r/Seedance_2_API