{"repo":"vitoplantamura/OnnxStream","free":true,"listed":false,"github":"https://github.com/vitoplantamura/OnnxStream","clone":"git clone https://github.com/vitoplantamura/OnnxStream.git","description":"Lightweight inference library for ONNX files, written in C++. It can run Stable Diffusion XL 1.0 on a RPI Zero 2 (or in 298MB of RAM) but also Mistral 7B on desktops and servers. ARM, x86, WASM, RISC-V supported. Accelerated by XNNPACK. Python, C# and JS(WASM) bindings available.","language":"C++","stars":2087,"topics":["machine-learning","onnx","raspberry-pi","stable-diffusion","llama","mistral","tinyml","wasm","webassembly","yolov8"],"license":null,"category":"machine-learning","readme_excerpt":"#### News 📣 - September 28, 2025: Added Python and C# bindings for the library! - April 15, 2025: Added WebAssembly demo of OpenAI's Whisper here (running in the browser). - September 19, 2024: Added WebAssembly support for the library! Demo of the YOLOv8 object detection model here (running in the browser). - January 14, 2024: Added LLM chat application ( TinyLlama 1.1B and Mistral 7B ) with initial GPU support! More info here. - December 14, 2023: Added support for Stable Diffusion XL Turbo 1.0 ! (thanks to @AeroX2) - October 3, 2023: Added support for Stable Diffusion XL 1.0 Base ! Index 👇 - Introduction - Stable Diffusion 1.5 - Stable Diffusion XL 1.0 Base - Stable Diffusion XL Turbo 1.0 - TinyLlama 1.1B and Mistral 7B - YOLOv8 (running in the browser) - OpenAI's Whisper (running in the browser) - Features of OnnxStream - Performance - Attention Slicing and Quantization - How OnnxStream Works - How to Build (Linux/Mac/Windows/Termux/FreeBSD) - How to Convert SD 1.5 Model - Related Projects - Credits OnnxStream The challenge is to run Stable Diffusion 1.5, which includes a large transformer model with almost 1 billion parameters, on a Raspberry Pi Zero 2, which is a microcomputer with 512MB of RAM, without adding more swap space and without offloading intermediate results on disk. The recommended minimum RAM/VRAM for Stable Diffusion 1.5 is typically 8GB. Generally major machine learning frameworks and libraries are focused on minimizing inference latency and/or maximizi","default_branch":null,"files":null,"tree":[],"storefront":"/r/vitoplantamura","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/vitoplantamura/OnnxStream/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}