HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllClosedOpen

Home / AI Models / Open

AI Models / Open 10 items

All Open items under AI Models on AIFIRST News, newest first. 10 items.

I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens

A developer hosted Kimi K3, Moonshot AI's 2.8 trillion parameter model, on 8 NVIDIA B300 GPUs via Modal, achieving 92 tokens per second at $190 per million tokens. Cold boot takes around 27 minutes to load the 1.56 TB model at a cost of roughly $56.79 per hour. The experiment shows frontier-scale inference is now accessible but remains expensive outside managed cloud services.

/24 Aug 2026/reddit

New Qwen 3.8 27B on a 39k line C to single-file HTML / Three.js port

A developer tested Qwen 3.8 27B against Opus 5 on a hard coding task: porting a 39,000-line procedural shooter written in C to a single-file HTML/Three.js web app. Qwen completed the port and produced a working result. The comparison highlights Qwen 3.8 27B's competitive coding ability for large-scale refactoring work.

/24 Aug 2026/reddit

We quantized Qwen 3.8 27B and compared the quants on an RTX 6000

A team created Atomic Dynamic GGUF quantizations of Qwen 3.8 27B and compared their performance on a voxel island creation task using an RTX 6000. Results showed strong quality retention across quantization levels. The post shares the quantized files and methodology for running Qwen 3.8 27B locally on consumer hardware.

/24 Aug 2026/reddit

You can now use MTP in GLM-Air

GLM-4.5-Air, a 106B mixture-of-experts model with only 12B active parameters, now supports Multi-Token Prediction in llama.cpp for a meaningful inference speed boost. The update makes it more practical for users who want a large MoE without requiring multiple high-end GPUs. MTP support is enabled via a configuration flag in the updated build.

/24 Aug 2026/reddit

Qwen 3.8 27B helped me with firmware and software preservation where Opus 4 couldn't

A developer used Qwen 3.8 27B to reverse-engineer and document firmware for old hardware, a task that Claude Opus 4 had not been able to complete. Qwen 3.8 27B analyzed the firmware binary, generated documentation, and helped write compatibility shims. The post highlights the model's practical utility for embedded systems and legacy software preservation work.

/24 Aug 2026/reddit

New 100B Liquid AI model coming soon

Liquid AI, known for fast Liquid Foundation Model (LFM) architectures and competitive small language models, is preparing a 100B parameter model release. This would be the company's largest model to date, following its track record of strong performance at smaller parameter counts.

/23 Aug 2026/reddit

Single RTX 5090: Qwen3.8-27B NVFP4 at 262K context in vLLM

A detailed setup guide documents running Qwen3.8-27B in NVFP4 quantization on a single RTX 5090 under vLLM, achieving 77 tokens/second at short context and 64.7 tokens/second at a real 262K token context window. The post provides reproducible configuration parameters for other 5090 owners targeting long-context local inference.

/23 Aug 2026/reddit

Palantir's Karp: frontier AI labs trying to drug addict us

Palantir CEO Alex Karp publicly accused frontier AI labs of building products designed to create addictive dependency rather than genuine utility, drawing comparisons to drug addiction dynamics. The statement comes amid broader industry debate about engagement-driven business models at OpenAI, Anthropic, and other major labs.

/23 Aug 2026/HN