HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllModelsInferenceToolingImage and Video
Open SourceInference 3 Sep 2026 HN

WebLLM: high-performance in-browser LLM inference engine

The brief

WebLLM enables high-performance large language model inference directly in the browser using WebGPU, with no server required.

Key points

  1. Models run fully client-side, making inference private and zero-cost after initial download.
  2. The project supports a range of open-weight models and is available as an open-source library.
Read the original

Sources

  • HNgithub.com