{"repo":"NVIDIA-AI-Blueprints/video-search-and-summarization","free":true,"listed":false,"github":"https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization","clone":"git clone https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git","description":"NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.","language":"C++","stars":1805,"topics":["rag","vlm","skills","video-analytics","video-search","computer-vision","generative-ai","long-video-understanding","model-context-protocol","multimodal-ai"],"license":null,"category":"ai-agents","readme_excerpt":"NVIDIA AI Blueprint: Video Search and Summarization (VSS) Build GPU-accelerated video AI agents that search, analyze, summarize, and reason over live or recorded video using natural language. NVIDIA AI Blueprint for Video Search and Summarization (VSS) combines vision-language models, RAG, and NVIDIA NIM microservices to deliver real-time video analytics, visual Q&A, alert verification, clip retrieval, and long-video summarization. - Search video streams or archives using natural language queries - Summarize hours of video - Ask visual questions and automatically generate reports - Detect and verify real-time alerts with VLMs 🚀 Try the Demo · ⚡ Quickstart · 📚 Documentation · 🏗️ Architecture · 📦 Latest Release Table of Contents - Overview - Use Case / Problem Description - Agent Workflows - Software Components - Target Audience - Repository Structure Overview - Documentation - Prerequisites - Hardware Requirements - Quickstart Guide - Contributing - License Overview The NVIDIA Blueprint for Video Search and Summarization (VSS) provides a suite of reference architectures for building vision agents and AI-powered video analytics applications. Those architectures bring together accelerated vision microservices, vision language models (VLMs), and large language models (LLMs) so you can use them in existing applications, as standalone microservices, or as part of a larger vision agent. VSS is organized into three areas of processing and analysis: real-time video intelligence (f","default_branch":null,"files":null,"tree":[],"storefront":"/r/NVIDIA-AI-Blueprints","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/NVIDIA-AI-Blueprints/video-search-and-summarization/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}