{"repo":"BrunoArsioli/llama-optimus","free":true,"listed":false,"github":"https://github.com/BrunoArsioli/llama-optimus","clone":"git clone https://github.com/BrunoArsioli/llama-optimus.git","description":"Lightweight Python tool using Optuna for tuning llama.cpp flags: towards optimal tok/s for your machine","language":"Python","stars":45,"topics":["automation","benchmark","llamacpp","localai","localllama","optimization","optuna"],"license":"MIT","category":"workflow-automation","readme_excerpt":"llama-optimus Running Local AI? llama-optimus will find the BEST llama.cpp performance flags for YOUR unique hardware llama-optimus is a lightweight Python tool to automatically optimize llama.cpp performance flags for maximum throughput. Maximize your tokens/s for prompt processing (pp) & token generation (tg). Brings Bayesian optimization (Optuna) to your local & embedded AI models. --- What does llama-optimus do? - Tunes llama.cpp parameters for maximum tokens/sec , using automated parameter search. - Bayesian optimization (Optuna) is used to maximize tokens/sec for prompt processing, generation or both - Estimates user-specified GPU layer count ( -ngl ). - Supports override patterns for --override-tensor : allows you to optimize advanced memory offloading for large models or low VRAM systems. - Built in system warmup to ensure benchmarking is done under real-world, “steady-state” conditions. - Grid search over categorical parameters (for flags like --override-tensor and --flash-attn ) combined with Bayesian tuning of numerical ones. - CLI interface: All major parameters and paths are settable via command line or environment variable. - Built on: Optuna for hyperparameter optimization and llama.cpp for inference. - Adapts to Apple Silicon, Linux x86, and NVIDIA GPU systems - Outputs copy-paste-ready commands for llama-server and llama-bench - Automatically runs llama-bench (at the end) comparing optimized vs. non-optimized results. - Quick & robust benchmarks: Controllable","default_branch":null,"files":null,"tree":[],"storefront":"/r/BrunoArsioli","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/BrunoArsioli/llama-optimus/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}