Evaluator
An evaluator turns a model's output into metrics. The LangiumEvaluator does this by running LLM output through your language, parsing it, validating it, and performing measurements on the results as a whole. This result is then computed into a grade, or score, that is used as a means of evaluating the quality of the model's output.
This page covers the evaluator objects and the evaluation matrix built on them. For the vitest-style API you use to write .eval.ts files, see Evals.
INFO
As a note, an evaluator grades outputs that are used to assess quality. It's important to point out that this isn't testing in the pass/fail sense.
Although it is technically testing, not all tests fall into this category.
This is tricky since, inherently, these checks are not quantitative, but we can reason about their quantities. For example, one parser error is a red flag, but what about a warning from some diagnostic? In the same sense, what about 50 warnings and one parser error? Do you account for ease of correction on errors, or whether multiple warnings signal to something greater that's afoot?
It really depends circumstantially, and so although this is a powerful way to assess your model's capabilities, it's not the only way to assess a model's capabilities in a standalone sense. It's always good to have the means to verify your model's performance in more than one way.
LangiumEvaluator
There are several exports you can leverage up front. These are the default Evaluator, the LangiumEvaluator, and a helper mergeEvaluators to sequence evaluators. The base Evaluator is extended to produce your own custom evaluators. The LangiumEvaluator provides some prebuilt handling to perform validations, and return diagnostics of note.
And here are the imports for reference:
import { Evaluator, LangiumEvaluator, mergeEvaluators } from 'langium-ai-tools/evaluator';To test this out, construct an evaluator using your language's services and hand it a string:
import { EmptyFileSystem } from 'langium';
import { LangiumEvaluator } from 'langium-ai-tools/evaluator';
// change for your own language
import { createMiniLogoServices } from "langium-minilogo/module";
import { LangiumDocumentAnalyzer } from "langium-ai-tools/analyzer";
const services = createMiniLogoServices(EmptyFileSystem).MiniLogo;
const evaluator = new LangiumEvaluator(services);
const result = await evaluator.evaluate(modelOutput, expectedOutput);
console.log(result.data.errors, result.data.warnings);evaluate is async, so make sure to await it! Additionally, EmptyFileSystem is the right choice when you're grading a string in memory, not opening files. You can also use the NodeFileSystem as well if you're doing checks on a workspace at large.
What's in a grade?
LangiumEvaluator measures whether the output is a valid program in your language. It parses the text, builds the document with validation enabled, and counts the diagnostics your own validators produce. It does not explicitly compare the output to an expected answer; instead, it gives you insight into the quantity of diagnostics emitted.
If you want similarity, that's better handled by a custom evaluator.
That distinction is the reason this evaluator is a good default: it's the one metric that correlates directly with your LSP diagnostics. These are already used, to great effect, to evaluate and guide an LLM's code generation attempts. In this case, you get a chance to view correctness without giving the model access to those same diagnostics.
Result shape
type EvaluatorResult<T> = {
name: string; // the evaluator's class name
metadata: { duration: number } & Record<string, unknown>;
data: T;
};It's worth noting that T for the LangiumEvaluator is bound to LangiumEvaluatorResultData.
As for the metrics that we get back, the numbers live in result.data:
| Field | Meaning |
|---|---|
errors | Diagnostics with severity 1. Anything greater than 0 means the output isn't syntactically or semantically valid. |
warnings | Severity 2. |
infos | Severity 3. |
hints | Severity 4. |
unassigned | Diagnostics that carried no severity. |
failures | 1 when the document couldn't be built at all (an exception during parse/build), otherwise 0. This is a fallback for outright failures that potentially crashed or experienced undefined behavior. |
diagnostics | The raw Diagnostic[], so you can inspect messages and ranges rather than just counts. |
Reading the diagnostics directly is often the most useful part when you're iterating on a system prompt, as it tells you which rules were broken.
for (const d of result.data.diagnostics) {
console.log(`${d.severity}: ${d.message}`);
}
// error: Could not resolve reference to AbstractRule named 'MISSING'.TIP
It's often the case that effective DSL stacks for an AI feed back LSP diagnostics to support self-correction. This can save on initial context, rather than trying to account for every possible error case that needs to be avoided in a one-shot response.
Code blocks are extracted automatically
If the input contains a fenced code block, evaluate grades the first fenced block rather than the whole string. Models frequently wrap generated code in prose and fences (or can be nudged to), so this saves a processing step:
// both of these grade the same program
await evaluator.evaluate('entity Person { name: string }', expectedResponse);
await evaluator.evaluate('Sure! Here you go:\n```mydsl\nentity Person { name: string }\n```', expectedResponse);Two consequences worth knowing: the fence language tag is ignored, and if a response contains several blocks only the first is looked at.
For when you have more than one language
await evaluator.evaluate(input, 'mydsl');The second parameter is a file extension, but it won't come up very often. It's used to help build the in-memory document's URI with the correct extension, so Langium picks the right language services to parse with. Pass it when your services host several languages and you need to set one explicitly, otherwise it'll default to your language's first extension defined in LanguageMetaData.
Trying it without your own DSL
Every Langium install ships a real Langium language: the grammar language itself. It's helpful for experimenting with evaluators before you set up your own services, and it's what the repo's own example project uses.
import { EmptyFileSystem } from 'langium';
import { createLangiumGrammarServices } from 'langium/grammar';
import { LangiumEvaluator } from 'langium-ai-tools/evaluator';
const services = createLangiumGrammarServices(EmptyFileSystem).grammar;
const evaluator = new LangiumEvaluator(services);
const result = await evaluator.evaluate(`
grammar Broken
entry Greeting: 'hello' name=MISSING;
`, expectedResponse);
// result.data.errors === 1
// result.data.diagnostics[0].message ===
// "Could not resolve reference to AbstractRule named 'MISSING'."Custom evaluators
For cases where the regular LangiumEvaluator isn't sufficient on its own, or you need an entirely different evaluation approach, go for a custom evaluator. Extend Evaluator and return whatever data fields you like:
import { Evaluator, type EvaluatorResultData, type EvaluatorResult } from 'langium-ai-tools/evaluator';
// customized evaluator payload data
export interface EditDistanceData extends EvaluatorResultData {
edit_distance: number
}
export class EditDistanceEvaluator extends Evaluator {
async evaluate(response: string, expected_response: string): Promise<EvaluatorResult<EditDistanceData>> {
return {
name: "edit-distance-evaluator",
metadata: {
// adjust to however long this took to run,
// or zero if negligible
duration: 0
},
data: {
edit_distance: levenshtein(response.trim(), expected_response.trim())
}
};
}
}The abstract Evaluator requires you to implement evaluate with two arguments (response, expected_response) and a plain data object is returned. EvaluatorResultData is effectively Record<string, unknown>, so your metric names are yours, and they can be passed straight through to reports and to the averaging helpers, which aggregate every numeric field they find.
Some common patterns for custom evaluators include: string similarity, presence of specific AST node types or structure, phrase detection, LLM as a judge, or checking against the embedding distance of a reference answer.
mergeEvaluators
const combined = mergeEvaluators(new LangiumEvaluator(services), new EditDistanceEvaluator());Runs each evaluator in sequence on the same input and shallow-merges their results into one object. Later evaluators win on key collisions, which is important to keep in mind. So be sure your metrics have distinct names if that's the case.
Evaluation matrix
import { EvalMatrix } from 'langium-ai-tools/evaluator';A single evaluator grades one output. An EvalMatrix runs the cross product: every runner against every case, scored by every evaluator, and repeated num_runs times, with the results aggregated and written to disk.
This is a helpful tool for answering questions about comparative performance between models, their stacks, and the tasks at hand.
Runners
A runner is anything that turns a prompt into text. Really, anything:
import { type Runner, type Message } from 'langium-ai-tools/evaluator';
const myRunner: Runner = {
name: 'my-model',
runner: async (prompt: string, messages: Message[]) => {
const response = await myModel.generate(prompt, messages);
return response;
}
};Message is { role: 'user' | 'system' | 'assistant', content: string }, which is the lowest common denominator across all providers (local or otherwise).
Because the interface is quite thin, a runner can wrap a direct model call, a RAG pipeline with a vector lookup in front of it, a multi-step agent, or a canned response for a sanity check.
TIP
Runner names do have to be unique, otherwise run() rejects duplicates up front rather than producing an ambiguous report.
Cases
Short for an evaluation case. These are the input-output pairs that are assessed. At the core they have a prompt going in, and some expected response coming out.
import { type EvalCase } from 'langium-ai-tools/evaluator';
const testCase: EvalCase = {
name: 'hello-world-grammar',
prompt: 'Generate a Langium grammar for a simple hello world DSL',
expected_response: `grammar HelloWorld
entry Greeting: 'hello' name=ID;
terminal ID: /[_a-zA-Z][\\w_]*/;`,
// optional
history: [{ role: 'system', content: 'You are an expert in Langium grammars.' }],
tags: ['grammar', 'beginner']
};Note that although there's an expected_response, an evaluator does not need to honor it. It may very well be an evaluator that's interested in other aspects of the output, rather than how well it adhered to an expected answer.
| Field | Required | Meaning |
|---|---|---|
name | ✓ | Identifies the case in results and reports. |
prompt | ✓ | The input handed to each runner. |
expected_response | ✓ | The reference answer, passed to evaluators as their second argument. |
history | Prior messages (system/user/assistant) prepended for this case. | |
tags | Free-form labels for your own filtering and grouping. |
Cases can also live in YAML, which can be easier to maintain once you have more than a few:
# eval-cases.yaml
eval_cases:
- name: "hello-world-grammar"
prompt: "Generate a Langium grammar for a simple hello world DSL"
expected_response: |
grammar HelloWorld
entry Greeting: 'hello' name=ID;
terminal ID: /[_a-zA-Z][\w_]*/;
tags:
- "grammar"
- "beginner"import { loadFromYaml } from 'langium-ai-tools/evaluator';
import { readFileSync } from 'node:fs';
const cases = loadFromYaml(readFileSync('eval-cases.yaml', 'utf-8'));loadFromYaml takes the YAML string, not a path. It accepts either a top-level eval_cases: list or a single case object, and validates as it goes. If there's a missing prompt or a non-boolean flag, it'll throw with the offending field name.
Running the matrix
Once you have your evaluators, your runners, and your cases, you can proceed with running an EvalMatrix. The process is pretty simple, just pass each of the above items in, and supply some base configuration.
import { EvalMatrix, LangiumEvaluator } from 'langium-ai-tools/evaluator';
const matrix = new EvalMatrix({
config: {
name: 'Model Comparison',
description: 'Comparing models for DSL generation',
history_folder: '.eval-history',
num_runs: 3
},
runners: [baseRunner, ragRunner],
evaluators: [
{ name: 'Edit Distance', eval: new EditDistanceEvaluator() }
],
cases
});
const results = await matrix.run();| Config field | Meaning |
|---|---|
name | Descriptive name; also used in the report filename. |
description | Longer description, stored in the report. |
history_folder | Directory for timestamped JSON reports. Created if missing. |
num_runs | How many times to run each runner & case combination, results are averaged later. |
Evaluators are supplied as { name, eval } pairs — the name is what shows up in results, so it can be more descriptive than the class name itself.
run() logs its progress (runner, case, evaluator, run number) as it goes, then writes the report and returns a flat result list.
Results
The eval matrix returns an array of results, where each result has an associated runner, case, evaluator, and iteration:
{
"name": "stub-runner - hello-grammar - Length Ratio",
"metadata": {
"runner": "stub-runner",
"evaluator": "Length Ratio",
"testCase": { "name": "hello-grammar", "prompt": "…", "expected_response": "…" },
"actual_response": "grammar Hello\nentry Greeting: 'hello' name=ID;",
"duration": 0.412,
"run_count": 1
},
"data": { "length_ratio": 0.7 }
}metadata.duration is the runner's time in seconds (how long generation took), run_count is the 1-based iteration, and actual_response is the full text that was evaluated. The latter is helpful when a score looks potentially wrong, and you need to see what the model actually responded with.
Every run also writes <timestamp>-<name>.json into the specified history_folder. Each run contains the config, the date, the total runtime, and all results, plus a last.txt naming the most recent report.
Aggregating
Naturally, when you have the ability to perform more than one iteration on the same runner-case-evaluator product, you'll also want a way to aggregate across them.
import {
averageAcrossCases,
averageAcrossRunners,
loadReport,
loadLastResults
} from 'langium-ai-tools/evaluator';
// average the num_runs iterations of each runner-case-evaluator combination
const perCase = averageAcrossCases(results);
// collapse further, to one row per runner
const perRunner = averageAcrossRunners(results);
console.table(perRunner.map((r) => ({ name: r.name, ...r.data })));Both helpers average every numeric field they find and drop the non-numeric ones. The results are rounded to two decimals as well.
averageAcrossCases groups by the combined runner–case–evaluator name, and averageAcrossRunners reduces to one row per runner — the table you actually want when comparing candidates.
Loading prior runs
Reading past runs can be done with loadReport and loadLastResults:
const report = loadReport('.eval-history/2026-01-01T00-00-00-000Z-model-comparison.json');
const lastThree = loadLastResults('.eval-history', 3);loadLastResults(dir, take) sorts the directory's report filenames newest-first and returns the results from the take most recent. Note that take counts reports, not individual results.