- Powered by AI Foundry
Evaluate, validate, and trust your RAG systems at scale.
Gen AIR is Techment's evaluation framework for Retrieval-Augmented Generation systems. It combines Azure AI Foundry's automated metrics with custom dashboards and human validation to give you end-to-end confidence in your AI's accuracy, relevance, and grounded ness.
Business Challenges
Enterprises deploying custom RAG-based chat bots face deep evaluation gaps that are invisible until they reach production.
No way to measure RAG performance
Teams lack structured tools to assess whether their AI is retrieving the right content and generating accurate answers.
Hallucinations go undetected
Generative models can produce plausible but unsupported answers. Without grounded ness checks, these slip through.
Retrieval quality is an unknown
Even when answers seem correct, the underlying document retrieval may be fetching the wrong or incomplete context.
No test data for custom systems
Custom RAG systems have no ready-made evaluation datasets. Generating realistic test cases manually is slow and costly.
Human review is disconnected
QA teams review AI outputs in silos with no structured interface to compare expected vs. actual responses and log feedback.
Continuous improvement has no baseline
Without reproducible evaluation metrics, it’s impossible to know whether model or prompt changes improved performance.
- Our Solution
Gen AIR — a full-cycle RAG evaluation platform
Gen AIR automates dataset creation, runs multi-metric evaluations, and surfaces results in a custom dashboard — with human validation built in at every stage.
PHASE 01
Dataset Creation
Synthetic Q&A from your PDFs
PHASE 02
Automated Evaluation
Multi-metric GPT-4.1 scoring
PHASE 03
Human Validation
Streamlit SME review workflow
- SYNTHETIC DATASET GENERATION
Auto-generates realistic Q&A datasets
From your own PDFs using AI Foundry — no manual test authoring needed.
- MULTI-METRIC EVALUATION ENGINE
Scores across 5 key dimensions
Similarity, retrieval, completeness, relevance, and grounded ness using GPT-4.1 as judge.
- CUSTOM STREAMLIT DASHBOARD
One unified view
Overall performance, per-metric scores, and human review workflows in one purpose-built UI.
- HUMAN-IN-THE-LOOP VALIDATION
Structured QA review
QA analysts compare agent responses to expected answers, mark correctness, and submit feedback directly.
- EDGE & NEGATIVE QUERY TESTING
Adversarial stress-testing
Augments datasets with adversarial and out-of-scope queries to stress-test system limits before production.
- CONTINUOUS IMPROVEMENT LOOP
Versioned, reproducible runs
Making it easy to measure the impact of every model or prompt change over time.
- How It Works
Three-Phase Evaluation Pipeline
Gen AIR follows a structured three-step process — from generating test data to validating results with human reviewers.
Dataset Creation
Generate a synthetic Q&A dataset from your reference PDFs in AI Foundry. Configure task type, question style, sample size (50–1000), and output format. Export as JSONL, enrich with real responses, citations, and edge cases, then convert back for evaluation.
Automated Evaluation
Upload the dataset to AI Foundry's evaluation module. Select GPT-4.1 or any other as the judge model. Auto-map data fields to evaluator metrics. Run the evaluation and review the dashboard for token usage, scores, and per-evaluator breakdowns.
Human Validation
QA reviewers use the Stream lit dashboard to compare query, agent response, and expected answer side by side. Mark responses as correct or incorrect, add feedback, and build a trusted ground truth dataset for ongoing improvement.
Dataset Creation
Generate a synthetic Q&A dataset from your reference PDFs in AI Foundry. Configure task type, question style, sample size (50–1000), and output format. Export as JSONL, enrich with real responses, citations, and edge cases, then convert back for evaluation.
Automated Evaluation
Upload the dataset to AI Foundry's evaluation module. Select GPT-4.1 or any other as the judge model. Auto-map data fields to evaluator metrics. Run the evaluation and review the dashboard for token usage, scores, and per-evaluator breakdowns.
Human Validation
QA reviewers use the Stream lit dashboard to compare query, agent response, and expected answer side by side. Mark responses as correct or incorrect, add feedback, and build a trusted ground truth dataset for ongoing improvement.
- Benefits & Value Proposition
Proven Performance - Out of the Box
Zero manual test authoring
Get your evaluation pipeline running without writing a single test case by hand.
Catch hallucinations early
Surface unsupported AI claims during testing — not after a production incident.
Identify retrieval gaps before users do
Find silent retrieval failures before your clients or internal operators encounter them.
Build stakeholder trust
Mathematical proof of AI accuracy and grounded ness for leadership and compliance teams.
Continuous, measurable improvement
Track the impact of every model update and prompt change with versioned, reproducible evidence.
Enterprise-ready, any RAG stack
Compatible across any custom RAG architecture — not locked to a single LLM provider.
Technical Stack
| Category | Component |
|---|---|
| AI & Evaluation Platform |
|
| Dashboard & Review Interface |
|
| Evaluation Metrics |
|
Sample Evaluation Scores — Production System
Scores achieved on a real enterprise RAG deployment.
| Metric | Score |
|---|---|
| Groundedness | 0.91 ✓ |
| Retrieval Score | 0.87 ✓ |
| Relevance | 0.94 ✓ |
| Completeness | 0.85 ✓ |
| Similarity | 0.74 ⚠ |