Splitter
Generic text splitters can chunk a document at some character count, a line count, on blanks, or with a regex. For a DSL that's quite arbitrary: it can break declarations in half, separate constructs from their comments, and produce chunks that can't be parsed on their own.
Thankfully, the solution is pretty easy: reuse the internal parser and abstract syntax types that Langium builds for you.
To this end, the splitter uses your language's own parser instead. You also describe the nodes you're interested in via predicates over the AST, and you get back chunks that line up with real syntactic boundaries. This is what you want before embedding documents into a vector DB, feeding examples to a model, or building a repo-level map of a codebase.
The splitter is fairly straightforward, with three entrypoints, in increasing order of control:
| Function | Returns | Use when |
|---|---|---|
splitByNode | string[] | You want text chunks along node boundaries. |
splitByNodeToAst | AstNode[] | You want the nodes themselves and will handle text yourself. |
ProgramMapper | string[] | You want a condensed summary line per node, not its source. |
And their imports for reference:
import {
splitByNode,
splitByNodeToAst,
ProgramMapper
} from 'langium-ai-tools/splitter';splitByNode
function splitByNode(
document: string,
nodePredicates: Array<(node: AstNode) => boolean> | ((node: AstNode) => boolean),
services: LangiumServicesLike,
options?: SplitterOptions
): string[]
| Parameter | Meaning |
|---|---|
document | The DSL source to split, as a string. |
nodePredicates | One predicate, or an array of them. A node is picked up if any predicate matches. |
services | Your language's services, which is the same object you get from createMyDslServices(...).MyDsl. Used to parse the document. |
options.commentRuleNames | Comment terminal rule names to pull into each chunk. Defaults to ['ML_COMMENT', 'SL_COMMENT']. You can pass [] to exclude comments. |
Predicates run over every node in the AST, so (node) => node.$type === 'Func' chunks by whatever rule corresponds to the Func AST type. Using Langium's generated AST types, you can import type guards (e.g. isFunc(node)) to perform the same task.
By default, comments become part of the chunk they document. For each matched node the splitter looks for an attached comment node and, if it finds one, starts the chunk at the comment instead of the declaration. That's usually what's desired for retrieval, especially for semantic search. When it isn't, you can pass commentRuleNames: [] to keep comments out of the chunk.
TIP
The default comment rule names assume Langium's conventional hidden terminal names. If your grammar calls its hidden comment terminals something else, pass those names, otherwise comments are silently left out.
Empty and unparsable input does not produce an error. An empty document, or one with lexer or parser errors, produces [] (with the errors logged to the console) rather than throwing. If you're splitting model-generated text, an empty result indicates the output is empty or didn't parse. To distinguish between the two, it's recommended you first parse and validate your program text before splitting.
splitByNodeToAst
function splitByNodeToAst(
document: string,
nodePredicates: Array<(node: AstNode) => boolean> | ((node: AstNode) => boolean),
services: LangiumServicesLike
): AstNode[]
This function is identical in terms of its predicate matching, but instead returns the matching AST nodes. Reach for this when you're building your own serialization logic, custom filtering or post-processing on nodes, cross-reference handling logic, or really anything that doesn't fall into the splitByNode function's jurisdiction. This one has no options parameter, since you're taking the serialization matter into your own hands.
ProgramMapper
class ProgramMapper {
constructor(services: LangiumServicesLike, options: ProgramMapOptions) {}
map(document: string): string[] { return []; }
}Where splitByNode gives you the source of each match, the ProgramMapper gives you a way to supply your own mapping rules for specific nodes. The key use case for this is when you need to create a program or repo map, usually used as a compressed representation of a program or repo to be placed into context. The closest analogy is to importing a header file to get declarations without definitions or implementation details, it's often enough information to determine what's going to be most helpful for your task.
| Field | Type | Meaning |
|---|---|---|
mappingRules | MappingRule[] | Applied in order; each has a predicate and a map. |
MappingRule.predicate | (node: AstNode) => boolean | Which nodes this rule handles. |
MappingRule.map | (node: AstNode) => string | The text to emit for a matched node. |
The result is one string per matched node, in document order. This constructs something akin to an outline of the program. These maps are cheap to build, small enough to sit in a system prompt, and enough for a model to know what exists and ask for the rest as needed.