{"repo":"vcache-project/vCache","free":true,"listed":false,"github":"https://github.com/vcache-project/vCache","clone":"git clone https://github.com/vcache-project/vCache.git","description":"Reliable and Efficient Semantic Prompt Caching with vCache","language":"Python","stars":77,"topics":["chatbot","consistency","correctness","gpt","guarantees","llama","llm","machine-learning","memcache","openai"],"license":null,"category":"ai-agents","readme_excerpt":"Reliable and Efficient Semantic Prompt Caching vCache is the first semantic prompt cache that guarantees user-defined error rate bounds. Semantic caching reduces LLM latency and cost by returning cached model responses for semantically similar prompts (not just exact matches). vCache replaces static thresholds with online-learned, embedding-specific decision boundaries —no manual fine-tuning required. This enables reliable cached response reuse across any embedding model or workload. 💳 Cost & Latency Optimization Reduce LLM API Costs by up to 10x. Decrease latency by up to 100x. 💡 Verified Semantic Prompt Caching Set an error rate bound—vCache enforces it while maximizing cache hits. 🏢 System Agnostic Infrastructure vCache uses OpenAI by default for both LLM inference and embedding generation, but you can configure any other inference setup. Quick Install Install vCache in editable mode: Then, set your OpenAI key: Finally, use vCache in your Python code: How vCache Works vCache intelligently detects when a new prompt is semantically equivalent to a cached one, and adapts its decision boundaries based on your accuracy requirements. This lets it return cached model responses for semantically similar prompts (not just exact matches) reducing both inference latency and cost without sacrificing correctness. Please refer to the vCache paper for further details. System Integration Semantic caches sit between the application server and the LLM inference backend. Applications can r","default_branch":null,"files":null,"tree":[],"storefront":"/r/vcache-project","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/vcache-project/vCache/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}