News / Models
Models 9 items
All Models stories, newest first. Every card links to a TLDR page with the original source.
PSA: Opus 4.8 redefines the effort scale
System card data suggests 4.8 low effort matches 4.7 max on problem solving, while 4.8 medium burns more tokens than 4.7 high. Max mode scores slightly below x-high on SWE-bench, hinting overthinking degrades results.
Claude Opus 4.8 ships with a focus on honesty
Opus 4.8 arrives as a refined 4.7: fail-to-disclose rate drops from 19.7% to 3.7%, Dynamic Workflows bring parallel subagents to Claude Code as a research preview, and Fast mode offers 2.5x speed at a third of the cost.
Community finding: Opus 4.7 reasoning degrades in non-English languages
Users report Opus 4.7 shows significantly degraded reasoning when prompted in other languages: shallow responses, missed context and lost chain-of-thought depth. Opus 4.6 does not show the issue.
Goodbye ElevenLabs… your FREE AND LOCAL replacement has arrived
Voicebox is a local-first TTS app: voice clone from seconds of audio, 23 languages, 5 TTS engines, DAW-style timeline for podcast/conversation production, 100% on-device. Free alternative to ElevenLabs. Community: wor...
Opus 4.7 summarizes 110 threads about itself: planning up, token burn brutal
A meta-analysis of 110 release-day threads and 2,187 comments: Opus 4.7 is better at planning and orchestration, but token burn is heavy (one user hit 1.06B tokens in a day) and non-max effort brings laziness and.
GLM-OCR: 0.9B-parameter OCR model tops OmniDocBench, runs on Ollama
GLM-OCR handles complex tables, code-heavy documents, seals and multilingual layouts with a two-stage layout-then-recognition pipeline. At 0.9B parameters it runs locally via Ollama, vLLM or SGLang.
Pingu Unchained: uncensored 120B model for AI red-teaming
Pingu Unchained is a 120B uncensored model built for adversarial security testing: no content filters, 75+ security tool integrations via MCP, full token logging for accountability, and identity verification required.
"You can fine-tune 100+ open-source models without writing code"
LLaMA-Factory provides a unified web UI + CLI for fine-tuning LLMs and VLMs without writing custom training code. Supports LLaMA, Mistral, Qwen, DeepSeek, Gemma, Phi, Yi, and 90+ others. Positioned as a no-code altern...