A critique argues that the Artificial Analysis "Intelligence" benchmark is fundamentally flawed, using Qwen 3.8 27B as a case study where benchmark scores diverge sharply from coding task performance. The post includes side-by-side comparisons showing where high benchmark rankings do not reflect real-world output quality.
/23 Aug 2026/Rreddit
A machine-readable registry of more than 50 AI web crawlers, published as ai-crawlers.json by Crawl Census. Each entry carries robots.txt tokens, user-agent strings, crawler purpose, and compliance status, with block-rate data measured across more than 4,194 domains. The dataset separates training crawlers from answer-engine and retrieval crawlers run by the same operator, including OpenAI, Anthropic, and Google, and records which agents state that they disregard robots.txt. Licensed CC BY 4.0 with a live API.
/22 Aug 2026/GGitHub
Surya is an open-source document intelligence model under 1B parameters that handles OCR, layout analysis, and table extraction in one toolkit. It scores 83.3% on the olmocr benchmark - best-in-class under 3B params -...
/2 Jun 2026/r/r/LovingOpenSourceAI
Researchers measured a functional wellbeing proxy in models and engineered prompts that maximize it. Exposed models give warmer replies and end conversations less, while benchmark scores stay flat.
/4 May 2026/r/r/LLM
Alibaba's "Happy Horse" model (open source, expected) surpassed Seedance 2 on image-to-video leaderboard. Low-engagement post, possible promotional content (no audio support noted by commenter, community skeptical of ...
/16 Apr 2026/r/r/Seedance_2_API
UK AISI tested Claude Mythos Preview on CTFs and cyber-range sims: it solved 73% of expert CTF tasks and became the first model to complete a 32-step corporate network attack simulation end to end in 3 of 10 attempts.
/15 Apr 2026/r/r/unknown