Hacker News
WebLLM: high-performance in-browser LLM inference engine
WebLLM is a high-performance inference engine that leverages WebGPU to run large language models directly in web browsers without server-side processing. It offers full OpenAI API compatibility, including streaming and structured JSON generation. The engine supports various models like Llama 3, Phi 3, and Mistral, and integrates into web applications via NPM or CDN.