llms.txt generator#
Generates llms.txt and llms-full.txt at the docs root from docs.json
and the MDX files in this repo. Replaces the auto-generated versions Mintlify
ships so we can control exactly what lands at
https://upstash.com/docs/llms.txt and /llms-full.txt.
Layout#
Outputs are written to the docs root:
docs/llms.txtdocs/llms-full.txt
Scripts#
check is what the GitHub workflow runs on every PR. If it finds the files
stale, the workflow commits the regenerated output back to the PR branch and
fails the run so the reviewer notices.
How it works#
1. walkNavigation (generator.ts)#
A generator that walks the navigation.tabs tree in docs.json and yields
one item per node:
- Tab names are emitted as groups at
depth: 0. Nested groups increase the depth by one per level. pagesandgroupsare mutually exclusive at every level — the types enforce that.- A bare-string
openapi(used at the tab level, e.g. Developer API) is yielded with onlysource. The object form{source, directory}is yielded with both.
2. expandOpenApi (openapi.ts)#
Loads an OpenAPI spec and emits one entry per operation, mirroring how Mintlify resolves spec-driven pages:
- If the op has
x-mint.href, use it verbatim (with the leading/stripped). - Else if a
directoryis given, the path is<directory>/<tag>/<kebab-summary>. - Else the path is
api-reference/<tag>/<kebab-summary>(the fallback Mintlify uses when only a tab-level openapi string is set).
Duplicate slugs within a single spec are disambiguated with -1, -2, …
matching upstream.
3. build.ts#
Collects entries from three sources:
README.mdxat the docs root (Mintlify includes it even though it isn't indocs.json).walkNavigation— every MDX page and every OpenAPI operation.- An orphan-MDX scan — any
.mdxfile in the tree that isn't referenced fromdocs.json, since Mintlify still publishes them.
Entries are sorted alphabetically by site-relative path (without .md).
Two files are then written:
-
llms.txt— flat## Docslist, then a## OpenAPI Specsfooter for object-form openapi refs whose containing group sits at depth 2. The depth filter is an empirical match to upstream and excludes deeper specs (e.g. QStash REST API at depth 3). -
llms-full.txt— per-page blocks in this shape:For OpenAPI ops the body is a synthesized
<spec> <method> <api-path>\n<description>block.For MDX pages the body is the file body (frontmatter stripped) with a few transforms Mintlify applies at build time:
- bullet
-→* - horizontal rule
---→*** - opening code fences with a language →
lang theme=system ``` - internal link URLs
/foo→/docs/foo \&→&inside link URLs- trailing whitespace stripped from every line
- bullet
Known divergences from upstream#
llms.txt is byte-identical to https://upstash.com/docs/llms.txt.
llms-full.txt is structurally identical to upstream but its MDX bodies
still differ in roughly 53K of 87K lines, all from transformations that
require a real MDX/JSX parser:
- JSX child re-indentation. Mintlify indents content inside
<Tabs>,<Tab>,<Accordion>,<Steps>, etc. by 4 spaces per nesting level. We emit those bodies verbatim. - Markdown table column padding. Upstream pads every cell so column boundaries align vertically. We don't.
<img>attribute injection. Some image tags receivesrc=…URLs injected from external data (e.g. shields.io badges). We can't reproduce these without Mintlify's component resolution data.- Blank-line spacing around components occasionally differs as a knock-on effect of the above.
Closing the gap would mean integrating a Mintlify-compatible MDX
serializer (@mdx-js/mdx or remark-mdx + remark-stringify) with
component-aware rules. It's tracked as future work.
Editing rules#
- Don't edit
llms.txtorllms-full.txtby hand. CI will overwrite any manual edit. - To change the output, edit the source under
src/and re-runnpm run build. - When updating
docs.jsonor adding/removing MDX files, the workflow will regenerate the outputs automatically on the next PR push.