Gen AIR | RAG Evaluation Framework Powered by Azure AI Foundry | Techment

Evaluate, Validate, and Trust Your RAG Systems at Scale with Gen AIR Tool

Gen AIR is Techment's evaluation framework for Retrieval Augmented Generation systems. It combines Azure AI Foundry's automated metrics with custom dashboards and human validation to give you end to end confidence in your AI's accuracy, relevance, and groundedness.

  • Zero manual test authoring
  • Five evaluation metrics, GPT-4.1 as judge
  • Human validation built in at every stage
  • Enterprise ready, any RAG stack

Deep Evaluation Gaps That Stay Invisible Until Production

Enterprises deploying custom RAG based chatbots face deep evaluation gaps that are invisible until they reach production.

No Way to Measure RAG Performance

Teams lack structured tools to assess whether their AI is retrieving the right content and generating accurate answers.

Hallucinations Go Undetected

Generative models can produce plausible but unsupported answers. Without groundedness checks, these slip through.

Retrieval Quality Is an Unknown

Even when answers seem correct, the underlying document retrieval may be fetching the wrong or incomplete context.

No Test Data for Custom Systems

Custom RAG systems have no ready made evaluation datasets. Generating realistic test cases manually is slow and costly.

Human Review Is Disconnected

QA teams review AI outputs in silos with no structured interface to compare expected vs. actual responses and log feedback.

Continuous Improvement Has No Baseline

Without reproducible evaluation metrics, it's impossible to know whether model or prompt changes improved performance.

Gen AIR: A Full Cycle RAG Evaluation Platform

Gen AIR automates dataset creation, runs multi metric evaluations, and surfaces results in a custom dashboard, with human validation built in at every stage.

Synthetic Dataset Generation

Auto generates realistic Q&A datasets from your own PDFs using AI Foundry. No manual test authoring needed.

Multi Metric Evaluation Engine

Scores responses across similarity, retrieval, completeness, relevance, and groundedness using GPT-4.1 as judge.

Custom Streamlit Dashboard

A purpose built UI showing overall performance, per metric scores, and human review workflows in one place.

Human in the Loop Validation

QA analysts compare agent responses to expected answers, mark correctness, and submit feedback directly in the tool.

Edge & Negative Query Testing

Augments datasets with adversarial and out of scope queries to stress test system limits before production.

Continuous Improvement Loop

Evaluation runs are reproducible and versioned, making it easy to measure the impact of every model or prompt change.

A Three Phase Evaluation Pipeline

Gen AIR follows a structured three step process, from generating test data to validating results with human reviewers.

Dataset Creation

Generate a synthetic Q&A dataset from your reference PDFs in AI Foundry. Configure task type, question style, sample size (50–1000), and output format. Export as JSONL, enrich with real responses, citations, and edge cases, then convert back for evaluation.

Automated Evaluation

Upload the dataset to AI Foundry's evaluation module. Select GPT-4.1 or any other as the judge model. Auto map data fields to evaluator metrics. Run the evaluation and review the dashboard for token usage, scores, and per evaluator breakdowns.

Human Validation

QA reviewers use the Streamlit dashboard to compare query, agent response, and expected answer side by side. Mark responses as correct or incorrect, add feedback, and build a trusted ground truth dataset for ongoing improvement.

Proven Performance, Out of the Box

Gen AIR's evaluation framework has been validated on real enterprise RAG deployments.

Zero manual test authoring
Catch hallucinations early
Identify retrieval gaps before users do
Build stakeholder trust
Continuous, measurable improvement
Enterprise ready, any RAG stack

Built on Azure AI Foundry & Purpose Built Review Tooling

AI & Evaluation Platform

  • Azure AI Foundry
  • GPT-4.1 (judge model)
  • AI Foundry Evaluation SDK
  • Synthetic Data Generation

Dashboard & Review Interface

  • Streamlit
  • AI Foundry SDK
  • Custom Human Review UI

Evaluation Metrics

  • Similarity
  • Retrieval score
  • Completeness
  • Relevance
  • Groundedness

Gen AIR FAQs

What is Gen AIR?
Gen AIR is Techment's evaluation framework for Retrieval Augmented Generation (RAG) systems. It combines Azure AI Foundry's automated metrics with custom dashboards and human validation to give end to end confidence in your AI's accuracy, relevance, and groundedness.
How does Gen AIR detect hallucinations?
Every response is scored for groundedness using GPT-4.1 as a judge model, alongside similarity, retrieval, completeness, and relevance. Human reviewers then validate results in the dashboard, so plausible but unsupported answers are caught before they reach production.
Do I need to create test data manually?
No. Gen AIR auto generates realistic Q&A datasets from your own PDFs using AI Foundry, with configurable task type, question style, sample size from 50 to 1000, and output format. Datasets can also be augmented with adversarial and out of scope queries for stress testing.
Which evaluation metrics does Gen AIR measure?
Gen AIR scores responses across five metrics: similarity, retrieval score, completeness, relevance, and groundedness, using GPT-4.1 or any other model as the judge.
How does human validation work in Gen AIR?
QA reviewers use a custom Streamlit dashboard to compare the query, agent response, and expected answer side by side. They mark responses as correct or incorrect, add feedback, and build a trusted ground truth dataset for ongoing improvement.
Can Gen AIR work with my existing RAG stack?
Yes. Gen AIR is enterprise ready and works with any RAG stack. Evaluation runs are reproducible and versioned, so you can measure the impact of every model or prompt change on any system.

Know Exactly How Your RAG System Performs, Before Your Users Do

Book a personalised demo and see Gen AIR evaluate a real RAG deployment end to end.

Request a Demo

Hello popup window