Evaluate, validate, and trust your RAG systems at scale.

Gen AIR is Techment's evaluation framework for Retrieval-Augmented Generation systems. It combines Azure AI Foundry's automated metrics with custom dashboards and human validation to give you end-to-end confidence in your AI's accuracy, relevance, and grounded ness.

Business Challenges

Enterprises deploying custom RAG-based chat bots face deep evaluation gaps that are invisible until they reach production.
No way to measure RAG performance

Teams lack structured tools to assess whether their AI is retrieving the right content and generating accurate answers.

Hallucinations go undetected

Generative models can produce plausible but unsupported answers. Without grounded ness checks, these slip through.

Retrieval quality is an unknown

Even when answers seem correct, the underlying document retrieval may be fetching the wrong or incomplete context.

No test data for custom systems

Custom RAG systems have no ready-made evaluation datasets. Generating realistic test cases manually is slow and costly.

Human review is disconnected

QA teams review AI outputs in silos with no structured interface to compare expected vs. actual responses and log feedback.

Continuous improvement has no baseline

Without reproducible evaluation metrics, it’s impossible to know whether model or prompt changes improved performance.

Gen AIR — a full-cycle RAG evaluation platform

Gen AIR automates dataset creation, runs multi-metric evaluations, and surfaces results in a custom dashboard — with human validation built in at every stage.
PHASE 01
Dataset Creation

Synthetic Q&A from your PDFs

PHASE 02
Automated Evaluation

Multi-metric GPT-4.1 scoring

PHASE 03
Human Validation

Streamlit SME review workflow

Auto-generates realistic Q&A datasets

From your own PDFs using AI Foundry — no manual test authoring needed.

Scores across 5 key dimensions

Similarity, retrieval, completeness, relevance, and grounded ness using GPT-4.1 as judge.

One unified view

Overall performance, per-metric scores, and human review workflows in one purpose-built UI.

Structured QA review

QA analysts compare agent responses to expected answers, mark correctness, and submit feedback directly.

Adversarial stress-testing

Augments datasets with adversarial and out-of-scope queries to stress-test system limits before production.

Versioned, reproducible runs

Making it easy to measure the impact of every model or prompt change over time.

Three-Phase Evaluation Pipeline

Gen AIR follows a structured three-step process — from generating test data to validating results with human reviewers.

Dataset Creation

Generate a synthetic Q&A dataset from your reference PDFs in AI Foundry. Configure task type, question style, sample size (50–1000), and output format. Export as JSONL, enrich with real responses, citations, and edge cases, then convert back for evaluation.

Automated Evaluation

Upload the dataset to AI Foundry's evaluation module. Select GPT-4.1 or any other as the judge model. Auto-map data fields to evaluator metrics. Run the evaluation and review the dashboard for token usage, scores, and per-evaluator breakdowns.

Human Validation

QA reviewers use the Stream lit dashboard to compare query, agent response, and expected answer side by side. Mark responses as correct or incorrect, add feedback, and build a trusted ground truth dataset for ongoing improvement.

Dataset Creation

Generate a synthetic Q&A dataset from your reference PDFs in AI Foundry. Configure task type, question style, sample size (50–1000), and output format. Export as JSONL, enrich with real responses, citations, and edge cases, then convert back for evaluation.

Automated Evaluation

Upload the dataset to AI Foundry's evaluation module. Select GPT-4.1 or any other as the judge model. Auto-map data fields to evaluator metrics. Run the evaluation and review the dashboard for token usage, scores, and per-evaluator breakdowns.

Human Validation

QA reviewers use the Stream lit dashboard to compare query, agent response, and expected answer side by side. Mark responses as correct or incorrect, add feedback, and build a trusted ground truth dataset for ongoing improvement.

Proven Performance - Out of the Box

Zero manual test authoring

Get your evaluation pipeline running without writing a single test case by hand.

Catch hallucinations early

Surface unsupported AI claims during testing — not after a production incident.

Identify retrieval gaps before users do

Find silent retrieval failures before your clients or internal operators encounter them.

Build stakeholder trust

Mathematical proof of AI accuracy and grounded ness for leadership and compliance teams.

Continuous, measurable improvement

Track the impact of every model update and prompt change with versioned, reproducible evidence.

Enterprise-ready, any RAG stack

Compatible across any custom RAG architecture — not locked to a single LLM provider.

Technical Stack

Category Component
AI & Evaluation Platform
  • Azure AI Foundry
  • GPT-4.1 (judge model)
  • AI Foundry Evaluation SDK
  • Synthetic Data Generation
Dashboard & Review Interface
  • Streamlit
  • AI Foundry SDK
  • Custom Human Review UI
Evaluation Metrics
  • Similarity
  • Retrieval Score
  • Completeness
  • Relevance
  • Groundedness

Sample Evaluation Scores — Production System

Scores achieved on a real enterprise RAG deployment.
Metric Score
Groundedness 0.91
Retrieval Score 0.87
Relevance 0.94
Completeness 0.85
Similarity 0.74

Ready to evaluate your RAG system with confidence?

Talk to Techment's AI team — get a tailored Gen AIR demo for your use case.

Hello popup window