Skip to main content

AI ContextRetrievalChunker

Text chunking utility for context retrieval in RAG (Retrieval-Augmented Generation) pipelines.

Splits Markdown documents, plain text or text extracted from PDF into controlled-size chunks with configurable overlap, preserving the document structure for better semantic retrieval quality.

Key features:

  • Structural segmentation: headings, paragraphs, code blocks, tables, lists and quotes are recognised and never split arbitrarily
  • Hierarchical context: every block carries the full heading tree (# Guide > ## Install > ### Windows)
  • Budget in characters or in real tokens (BPE cl100k_base / o200k_base)
  • Code blocks larger than the budget are split by lines with the fence reopened in each part
  • Tables larger than the budget repeat the header row in each part
  • Sentence-aware overlap, never in the middle of a word
  • No text loss: every block of the document belongs to exactly one chunk
  • Deterministic: the same input always produces the same chunks and the same id values, which makes re-ingestion idempotent
// Basic example
const chunker = _ai.contextRetrievalChunker()
const chunks = chunker.markdown(markdownDocument)

for (const chunk of chunks.listOfValues()) {
_log.info(`Chunk ${chunk.getInt('index')}: ${chunk.getString('breadcrumb')}`)
_log.info(`Text: ${chunk.getString('text')}`)
}

// Budget in real tokens and ingestion into the vector store
const client = _ai.client()
const vector = _ai.vector('default')

const blocks = _ai.contextRetrievalChunker()
.unit('tokens')
.chunkSize(320)
.overlap(48)
.source('manual-v1')
.markdown(markdownDocument)

for (const block of blocks.listOfValues()) {
const response = client.embeddings('embeddinggemma:latest', block.getString('text'))
const embedding = response.getValues('data').getValues(0).getValues('embedding')
vector.add('netuno', block.getString('id'), embedding, block.getString('text'), block.getValues('metadata'))
}


chunk


chunk(content: string) : Values

Description

Splits a document into chunks, automatically detecting whether the content is Markdown or running text. This is the entry point to use when the content origin is not known upfront.

How To Use
const chunks = chunker.chunk(content)

for (const chunk of chunks.listOfValues()) {
_log.info(chunk.getString('text'))
}
Attributes
NAMETYPEDESCRIPTION
contentstringText to split into chunks.
Return

( Values )

List of chunks. See the field descriptions in markdown.


chunk(content: string, chunkSize: int, overlap: int) : Values

Attributes
NAMETYPEDESCRIPTION
contentstring
chunkSizeint
overlapint
Return

( Values )


chunk(content: string, options: Values) : Values

Description

Splits a document into chunks with options, automatically detecting whether the content is Markdown or running text.

How To Use
const chunks = chunker.chunk(content, _val.map()
.set('unit', 'tokens')
.set('chunkSize', 320)
.set('overlap', 48))
Attributes
NAMETYPEDESCRIPTION
contentstringText to split into chunks.
optionsValuesOptions overriding the instance configuration. See the full list in markdown.
Return

( Values )

List of chunks. See the field descriptions in markdown.


chunkSize


chunkSize(chunkSize: int) : ContextRetrievalChunker

Description

Sets the maximum size of each chunk, in the unit chosen with unit. Default value: 1024 characters, or 256 when the unit is tokens and the size was not set explicitly.

How To Use
const chunks = chunker.chunkSize(1500).markdown(document)
Attributes
NAMETYPEDESCRIPTION
chunkSizeintMaximum size of each chunk, clamped to [32, 200000].
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


contextualize


contextualize(client: Client, document: string, chunks: Values) : Values

Description

Fills the context field of each chunk with a retrieval note, generated by the model from the whole document, and rebuilds the text field with that note included. This is the contextual retrieval technique: a chunk that only says "the value went up 3%" also comes to say which company and which quarter it is about, which greatly reduces retrieval misses.

The note has two parts, because a query arrives in two shapes. Two or three sentences situate the chunk by naming the product, the task and the conditions it takes for granted, which answers whoever describes the problem in their own words. Then a line of terms gathers the exact vocabulary: names, identifiers, configuration keys, commands, error codes, acronyms with what they stand for, synonyms and, on documents not written in English, the English term next to the original. That is what finds whoever pastes an error message or the name of a parameter.

It is a paid operation: it makes one model call per chunk. Run it during ingestion only, never per request.

The calls are sequential on purpose, because Client keeps session and token accounting state that is not safe to share across parallel calls. An error on one chunk does not stop the rest: it is logged, that chunk's context stays empty and processing continues.

Accepted options: model, temperature (0 by default, so the result is reproducible), documentMaxChars (truncates very large documents), template (placeholders {document}, {chunk}, {breadcrumb} and {heading}), system (system message), skipIfPresent (does not redo chunks that already have context) and failFast.

How To Use
const client = _ai.client()
const chunks = chunker.markdown(document)

chunker.contextualize(client, document, chunks)

for (const chunk of chunks.listOfValues()) {
_log.info(chunk.getString('context'))
// text already includes the context, it is what should be embedded
}
Attributes
NAMETYPEDESCRIPTION
clientClientAI client used to generate the context.
documentstringThe whole document, the same one the chunks came from.
chunksValuesList of chunks returned by markdown, text, pdf or chunk.
Return

( Values )

The same list of chunks, with context and text updated.


contextualize(client: Client, document: string, chunks: Values, options: Values) : Values

Description

Fills the context field of each chunk with options. See the full description and option list in contextualize.

How To Use
chunker.contextualize(client, document, chunks, _val.map()
.set('model', 'gpt-4o-mini')
.set('documentMaxChars', 40000))
Attributes
NAMETYPEDESCRIPTION
clientClientAI client used to generate the context.
documentstringThe whole document, the same one the chunks came from.
chunksValuesList of chunks returned by markdown, text, pdf or chunk.
optionsValuesContext generation options.
Return

( Values )

The same list of chunks, with context and text updated.


countTokens


countTokens(text: string) : int

Description

Counts the tokens of a text with the encoding configured in encoding, using the same BPE algorithm the models use. Useful to budget prompts and to check that a chunk fits the embedding model limit.

How To Use
_log.info('Tokens: '+ chunker.countTokens(text))
Attributes
NAMETYPEDESCRIPTION
textstringText to measure.
Return

( int )

Number of tokens in the text.


countTokens(text: string, encoding: string) : int

Description

Counts the tokens of a text with a specific encoding, or with the encoding inferred from a model name.

Attributes
NAMETYPEDESCRIPTION
textstringText to measure.
encodingstringEncoding name, such as cl100k_base or o200k_base, or a model name.
Return

( int )

Number of tokens in the text.


embed


embed(embed: string) : ContextRetrievalChunker

Description

Sets what goes into the text field, which is always the string meant to be embedded. The content field always keeps the real chunk and is what a search should return.

  • full, the default: context header, generated context and chunk body
  • context: header and generated context only, leaving the body out
  • content: body only, with no header and no context

The context mode indexes the generated note instead of the chunk. The note is dense prose, while the body carries code blocks, table pipes and markup that dilute the vector. In exchange, an exact term that exists in the body but not in the note stops being searchable, so this mode only makes sense after running contextualize. While the context is still empty the body is used anyway, so a heading is never indexed on its own.

How To Use
// Index the context, return the real text
const chunks = chunker.embed('context').markdown(document)
chunker.contextualize(client, document, chunks)

for (const chunk of chunks.listOfValues()) {
const response = client.embeddings(model, chunk.getString('text'))
const embedding = response.getValues('data').getValues(0).getValues('embedding')
vector.add('docs', chunk.getString('id'), embedding, chunk.getString('content'), metadata)
}
Attributes
NAMETYPEDESCRIPTION
embedstringfull, context or content.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


encoding


encoding(encoding: string) : ContextRetrievalChunker

Description

Sets the encoding used for token counting: cl100k_base (default), o200k_base, p50k_base, r50k_base, or the name of an OpenAI model, from which the encoding is inferred.

Attributes
NAMETYPEDESCRIPTION
encodingstringEncoding or model name.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


getChunkSize


getChunkSize() : int

Return

( int )


getEmbed


getEmbed() : string

Return

( string )


getEncoding


getEncoding() : string

Return

( string )


getMetadata


getMetadata() : Values

Return

( Values )


getMinChunkSize


getMinChunkSize() : int

Return

( int )


getOverlap


getOverlap() : int

Return

( int )


getSource


getSource() : string

Return

( string )


getUnit


getUnit() : string

Return

( string )


headingPath


headingPath(headingPath: boolean) : ContextRetrievalChunker

Description

Sets whether the context header uses the full heading tree, enabled by default, or only the nearest heading. The full tree gives much better retrieval on documents with nested sections, because a chunk under ### Windows also keeps # Guide and ## Install.

Attributes
NAMETYPEDESCRIPTION
headingPathbooleanUse the full tree or only the nearest heading.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


isHeadingPath


isHeadingPath() : boolean

Return

( boolean )


isPrependHeading


isPrependHeading() : boolean

Return

( boolean )


isSplitOnHeadings


isSplitOnHeadings() : boolean

Return

( boolean )


isStripDataUri


isStripDataUri() : boolean

Return

( boolean )


isStripHtmlComments


isStripHtmlComments() : boolean

Return

( boolean )


markdown


markdown(markdown: string) : Values

Description

Splits a Markdown document into chunks, using the instance configuration.

The document is first segmented into atomic blocks, respecting the markup: headings, paragraphs, code blocks, tables, lists and quotes. A # comment inside a code block is never mistaken for a heading, because detection is done with markup state. The blocks are then packed up to the budget, preferring to start at a heading.

Every chunk carries the heading tree in effect, prepended to the text field, which substantially improves semantic retrieval on documents with nested sections.

How To Use
const chunks = chunker.markdown('# Title\n\nDocument content...')

for (const chunk of chunks.listOfValues()) {
_log.info(chunk.getString('breadcrumb') +' -> '+ chunk.getInt('tokens') +' tokens')
_log.info(chunk.getString('text'))
}
Attributes
NAMETYPEDESCRIPTION
markdownstringText in Markdown format to split into chunks.
Return

( Values )

List of chunks, each with the fields: id (stable identifier, suitable for idempotent reindexing), hash (content digest), index and total (position and total), start and end (positions in the normalized text), length and tokens (size in characters and in tokens), heading and headingLevel (nearest heading and level), path (heading tree as a list), breadcrumb (the same tree as text), sections (every section the chunk touches), header (the rendered context header), content (chunk body), context (retrieval note, situating sentences plus terms, filled in by contextualize), text (the string to embed, composed according to embed), embed (the mode used), type (markdown, text or pdf), blocks (block types present), overlap (characters repeated from the previous chunk), page (the page the chunk starts on) and pages (every page the chunk spans), both only on paginated PDF, metadata (configured metadata) and synthetic (true when the body is not a literal slice of the source, because a code fence was reopened or a table header repeated).


markdown(markdown: string, chunkSize: int, overlap: int) : Values

Description

Splits a Markdown document into chunks with explicit size and overlap, in the unit configured with unit.

How To Use
// Chunks of 1500 characters with overlap of 200
const chunks = chunker.markdown(markdown, 1500, 200)
Attributes
NAMETYPEDESCRIPTION
markdownstringText in Markdown format to split into chunks.
chunkSizeintMaximum size of each chunk.
overlapintOverlap between consecutive chunks, clamped to half the chunk size.
Return

( Values )

List of chunks. See the field descriptions in markdown.


markdown(markdown: string, options: Values) : Values

Description

Splits a Markdown document into chunks with options, overriding the instance configuration for this call only.

Accepted options: chunkSize, overlap, minChunkSize, unit (chars or tokens), embed (full, context or content), encoding, prependHeading, headingPath, splitOnHeadings, stripDataUri, stripHtmlComments, source and metadata.

How To Use
const chunks = chunker.markdown(markdown, _val.map()
.set('unit', 'tokens')
.set('chunkSize', 320)
.set('overlap', 48)
.set('source', 'manual-v1')
.set('metadata', _val.map().set('language', 'en')))
Attributes
NAMETYPEDESCRIPTION
markdownstringText in Markdown format to split into chunks.
optionsValuesOptions overriding the instance configuration.
Return

( Values )

List of chunks. See the field descriptions in markdown.


metadata


metadata(metadata: Values) : ContextRetrievalChunker

Description

Sets metadata applied to every chunk, copied into the metadata field and ready to pass straight to vector.add.

How To Use
const chunks = chunker
.metadata(_val.map().set('source', 'manual').set('version', 3))
.markdown(document)
Attributes
NAMETYPEDESCRIPTION
metadataValuesMetadata applied to every chunk.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


minChunkSize


minChunkSize(minChunkSize: int) : ContextRetrievalChunker

Description

Sets the minimum size of a chunk. Chunks below this value are absorbed by a neighbour, to avoid orphan micro-chunks that pollute the vector store. Default value: a quarter of the chunk size.

Attributes
NAMETYPEDESCRIPTION
minChunkSizeintMinimum size of a chunk.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


overlap


overlap(overlap: int) : ContextRetrievalChunker

Description

Sets the overlap between consecutive chunks, in the unit chosen with unit. The overlap is built from the trailing sentences of the previous chunk, never mid-word, and is clamped to half the chunk size. Default value: 128 characters.

How To Use
const chunks = chunker.chunkSize(1500).overlap(200).markdown(document)
Attributes
NAMETYPEDESCRIPTION
overlapintOverlap between consecutive chunks, clamped to half the chunk size.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


pdf


pdf(pdfText: string) : Values

Description

Splits text extracted from a PDF into chunks, first applying the cleanup specific to this source: joining words hyphenated at the end of a line, removing headers and footers repeated across pages, and removing lines that are just the page number. When the text carries page separators, each chunk also gets the page field.

How To Use
const text = _pdf.toText(_storage.filesystem('server', 'docs', 'manual.pdf'))
const chunks = chunker.source('manual.pdf').pdf(text)

for (const chunk of chunks.listOfValues()) {
_log.info('Page '+ chunk.getInt('page') +': '+ chunk.getString('content'))
}
Attributes
NAMETYPEDESCRIPTION
pdfTextstringText extracted from a PDF.
Return

( Values )

List of chunks. See the field descriptions in markdown.


pdf(pdfText: string, chunkSize: int, overlap: int) : Values

Attributes
NAMETYPEDESCRIPTION
pdfTextstring
chunkSizeint
overlapint
Return

( Values )


pdf(pdfText: string, options: Values) : Values

Description

Splits text extracted from a PDF into chunks with options. See the accepted options in markdown.

Attributes
NAMETYPEDESCRIPTION
pdfTextstringText extracted from a PDF.
optionsValuesOptions overriding the instance configuration.
Return

( Values )

List of chunks. See the field descriptions in markdown.


prependHeading


prependHeading(prependHeading: boolean) : ContextRetrievalChunker

Description

Sets whether the context header is prepended to the text field of each chunk. Enabled by default. The content field always keeps the chunk body without the header.

Attributes
NAMETYPEDESCRIPTION
prependHeadingbooleanWhether to prepend the context header.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


source


source(source: string) : ContextRetrievalChunker

Description

Sets the document source identifier, used to build the id of each chunk. With the same source and the same content the id values are always the same, which allows reindexing a document without duplicating records in the vector store.

How To Use
const chunks = chunker.source('manual-v1').markdown(document)
// id -> manual-v1#0-3f2a1c9d
Attributes
NAMETYPEDESCRIPTION
sourcestringDocument source identifier.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


splitOnHeadings


splitOnHeadings(splitOnHeadings: boolean) : ContextRetrievalChunker

Description

Sets whether a new chunk preferably starts at a heading, enabled by default. This is what aligns chunks with the document sections.

Attributes
NAMETYPEDESCRIPTION
splitOnHeadingsbooleanWhether to preferably break at headings.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


stripDataUri


stripDataUri(stripDataUri: boolean) : ContextRetrievalChunker

Description

Sets whether images embedded as data: URIs are reduced to their content type, enabled by default. A single base64 PNG can take hundreds of thousands of characters with no semantic value at all.

Attributes
NAMETYPEDESCRIPTION
stripDataUribooleanWhether to reduce embedded images.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


stripHtmlComments


stripHtmlComments(stripHtmlComments: boolean) : ContextRetrievalChunker

Description

Sets whether HTML comments are removed from the Markdown, enabled by default. Comments inside code blocks are always preserved.

Attributes
NAMETYPEDESCRIPTION
stripHtmlCommentsbooleanWhether to remove HTML comments.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.


text


text(text: string) : Values

Description

Splits running text into chunks, packing whole paragraphs and cutting by sentences only when a paragraph goes over the budget. Numbered titles, in the 3.1 Install shape, are recognised as headings and feed the context tree.

How To Use
const chunks = chunker.text(plainText)
Attributes
NAMETYPEDESCRIPTION
textstringRunning text to split into chunks.
Return

( Values )

List of chunks. See the field descriptions in markdown.


text(text: string, chunkSize: int, overlap: int) : Values

Attributes
NAMETYPEDESCRIPTION
textstring
chunkSizeint
overlapint
Return

( Values )


text(text: string, options: Values) : Values

Description

Splits running text into chunks with options. See the accepted options in markdown.

Attributes
NAMETYPEDESCRIPTION
textstringRunning text to split into chunks.
optionsValuesOptions overriding the instance configuration.
Return

( Values )

List of chunks. See the field descriptions in markdown.


unit


unit(unit: string) : ContextRetrievalChunker

Description

Sets the budget unit: chars for characters or tokens for real tokens, counted with the same BPE algorithm the models use. Measuring in tokens is what guarantees no chunk goes over the embedding model limit.

When switching to tokens without having set chunkSize and overlap explicitly, the defaults become 256 and 32 tokens.

How To Use
const chunks = chunker.unit('tokens').chunkSize(320).markdown(document)
Attributes
NAMETYPEDESCRIPTION
unitstringchars or tokens.
Return

( ContextRetrievalChunker )

The instance itself, to chain configuration calls.