{"repo":"iaalm/llama-api-server","free":true,"listed":false,"github":"https://github.com/iaalm/llama-api-server","clone":"git clone https://github.com/iaalm/llama-api-server.git","description":"A OpenAI API compatible REST server for llama.","language":"Python","stars":206,"topics":["llama","llm","rest-api","language-model","openai","openai-api","privatization","selfhost"],"license":"MIT","category":"ai-agents","readme_excerpt":"🎭🦙 llama-api-server ======= This project is under active deployment. Breaking changes could be made any time. Llama as a Service! This project try to build a REST-ful API server compatible to OpenAI API using open source backends like llama/llama2. With this project, many common GPT tools/framework can compatible with your own model. 🚀Get started Try it online! Follow instruction in this collab notebook to play it online. Thanks anythingbutme for building it! Prepare model llama.cpp If you you don't have quantized llama.cpp, you need to follow instruction to prepare model. pyllama If you you don't have quantize pyllama, you need to follow instruction to prepare model. Install Use following script to download package from PyPI and generates model config file config.yml and security token file tokens.txt . Call with openai-python 🛣️Roadmap Tested with - [X] openai-python - [X] OPENAI\\ API\\ TYPE=default - [X] OPENAI\\ API\\ TYPE=azure - [X] llama-index Supported APIs - [X] Completions - [X] set temperature , top p , and top k - [X] set max tokens - [X] set echo - [ ] set stop - [ ] set stream - [ ] set n - [ ] set presence penalty and frequency penalty - [ ] set logit bias - [X] Embeddings - [X] batch process - [X] Chat - [ ] Prefix cache for chat - [ ] List model Supported backends - [X] llama.cpp via llamacpp-python - [X] llama via pyllama - [X] Without Quantization - [X] With Quantization - [X] Support LLAMA2 Others - [X] Performance parameters like n batch and n thread - [","default_branch":null,"files":null,"tree":[],"storefront":"/r/iaalm","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/iaalm/llama-api-server/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}