MARGINNOOK / DEVELOPER PREVIEW
Document paragraph search
Find full paragraphs matching literal keywords.
Developer preview. Hosted API is not live yet. Explore fixed synthetic samples; anonymous uploads, hosted execution and payments remain closed.
Inputs and results
Input: Text, HTML or ordinary DOCX. Returns: Matching paragraphs and Unicode coordinates in normalized text.
HTTP business ID document.search maps to MCP search_document.
Arguments
{
"query": "retry timeout",
"limit": 5,
"offset": 0
}query is required (512 bytes / 24 keywords). limit 1–10 (default 5). Matching is lexical FTS5 any-term, not semantic retrieval; no match does not prove absence. HTML/DOCX extraction may omit content.
Reproducible sample
Document paragraph search
Fixed synthetic sample · generated by the local processing core · no upload or live API call.
1. Upload input
{
"format": "text",
"content": "# Retries\n\nRetry a timeout with the same logical key.\n\n# Uploads\n\nDo not automatically repeat an uncertain upload."
}2. Submit a job
{
"tool": "document.search",
"resource_id": "<private-resource-handle>",
"arguments": {
"query": "retry timeout",
"limit": 5,
"offset": 0
}
}3. Operation output
This is the operation output inside result.output, not a complete job envelope.
{
"operation": "document.search",
"scope": {
"format": "text",
"indexed_sections": 4,
"retrieval": "lexical-fts5-any-term",
"terms": [
"retry",
"timeout"
],
"source_coordinates": "normalized_text",
"html_extraction_may_omit_content": false
},
"source_refs": [
"input:content"
],
"absence_proven": false,
"notice": "No keyword match does not prove absence. Results contain full matched paragraphs, not a semantic answer or the full document.",
"offset": 0,
"next_offset": null,
"total_items": 1,
"omitted_items": 0,
"incomplete": false,
"source_read_required": false,
"oversized_item_refs": [],
"sections": [
{
"section_id": "section-00002",
"text": "Retry a timeout with the same logical key.",
"heading": "# Retries",
"source_ref": {
"coordinate_space": "normalized_text",
"start_char": 11,
"end_char": 53
}
}
]
}Result fields and evidence
scope declares what was processed. incomplete, next_offset, omitted_items and source_read_required describe the bounded page. The core result is at most 128 KiB; whole items are omitted instead of clipped. A download is the complete returned page, not an unlimited result.
sections carry section_id, heading, text and start_char/end_char in normalized text. HTTP returns normalized_source_resource; character offsets are Unicode characters, not bytes.
Use GET /v1/resources/<private-resource-handle> for stored original/result bytes, or MCP read_evidence(resource_id, offset=0, length=4000). Never publish real handles. Resource-family expiry and management deletion apply to all derived evidence.
Errors and recovery
Quote checks arguments and metadata, not parsing. Preserve input and job ID on failure. Retry keyed submission with the original key; do not repeat uncertain uploads automatically. Invalid input needs correction; quota and platform capacity use distinct machine codes. Successfully extracting failing tests is a completed tool job.
HTTP contract · MCP · Limits · Switch samples