{"repo":"paradime-io/dbt-llm-evals","free":true,"listed":false,"github":"https://github.com/paradime-io/dbt-llm-evals","clone":"git clone https://github.com/paradime-io/dbt-llm-evals.git","description":"The warehouse-native LLM evaluation package for dbt™ - monitor AI quality without data egress","language":"Python","stars":30,"topics":["ai","bigquery","databricks","dbt","dbt-packages","llm-eval","snowflake","anthropic","gemini","llm"],"license":null,"category":"ai-agents","readme_excerpt":"Warehouse-native LLM evaluation and monitoring for dbt™ projects Quick Start Guide » Package Overview » Architecture Docs » Project Structure » Join the Discussion » ⭐️ Star the repo if this helps your LLM monitoring! A complete dbt™ package for evaluating LLM outputs directly within your data warehouse using warehouse-native AI functions. No external API calls, no data egress - everything runs inside your existing data infrastructure. What are LLM Evaluations? LLM evaluations (or \"evals\") are systematic methods for measuring the quality, accuracy, and performance of Large Language Model outputs. When deploying AI models in production, you need to continuously monitor whether your model is producing high-quality results that meet your business requirements. Why Evaluate LLM Outputs? - Quality Assurance : Ensure AI-generated content meets your standards - Performance Monitoring : Track model performance over time - Drift Detection : Identify when model outputs change unexpectedly - Business Confidence : Provide measurable metrics for AI system reliability LLM-as-a-Judge Framework This package uses the \"LLM-as-a-Judge\" approach, where another LLM evaluates the quality of your AI model's outputs. Instead of expensive human evaluation, a judge model: 1. Receives your original prompt, input data, and AI-generated output 2. Compares against baseline examples of good outputs 3. Evaluates across multiple criteria (accuracy, relevance, tone, etc.) 4. Scores each output on a 1-10 scale","default_branch":null,"files":null,"tree":[],"storefront":"/r/paradime-io","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/paradime-io/dbt-llm-evals/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}