Docs
Markdown pipeline
Follow a Markdown page through the Proa Docs build pass.
proa_docs_build compiles Markdown into escaped static HTML at build time, so no parser runs when a reader loads a page. This page follows that pass through MDX lowering, code fences, Mermaid, and search text.
How does the pipeline run?
One pass records heading IDs, table-of-contents data, links, search text, code blocks, file trees, Mermaid shells, optional configured MDX helper lowering, and optional type-table requests.
source .md or configured .mdx
-> optional MdxConfig lowering
-> pulldown-cmark parser
-> escaped HTML fragments
-> TOC entries from headings
-> link references for validation
-> search text from visible content and code
-> generated DocPage fields
This keeps runtime docs rendering static. The generated registry stores trusted HTML that the compiler produced, not user-provided HTML strings.
Each stage starts from what the parser recognizes.
What does the default parser enable?
The default parser enables:
- headings with generated IDs
- tables
- strikethrough
- task lists
- fenced code blocks
- links and images
- table-of-contents extraction
- search text extraction
The compiler escapes inline HTML unless it recognizes a configured type-table directive. It also escapes link and image URLs, and it strips javascript: and vbscript: URLs before writing HTML.
Headings get one more treatment: a stable id.
Generating heading IDs and TOC entries
Every Markdown heading gets a deterministic id:
## Render Traits
## Render Traits
The compiler emits:
<h2 id="render-traits">Render Traits</h2>
<h2 id="render-traits-2">Render Traits</h2>
It also records matching TocItem values:
pub struct TocItem {
pub depth: u8,
pub id: &'static str,
pub title: &'static str,
}
Those IDs power DocsToc, heading anchors, and link validation for #hash targets.
Authors rarely write raw Markdown alone, so the pass can lower a helper subset first.
Lowering configured MDX helpers
proa_docs_build does not execute arbitrary MDX. It can lower a small supported subset before Markdown parsing:
let markdown = proa_docs_build::MarkdownPipeline::default()
.mdx(proa_docs_build::MdxConfig::proa());
MdxConfig::proa() enables the helper subset this site uses:
| Authored helper | Lowered output |
|---|---|
<Cards> with <Card /> children | Markdown bullet links. |
<Files> with <Folder> and <File> children | A proa-filetree code fence that renders as a structured file tree. |
<AgentTerminal /> | A static text code fence. |
<img ...> | Markdown image syntax with normalized static paths. |
Layout <div> wrappers | The lowerer drops them. |
That keeps the runtime registry static. Downstream apps can start from MdxConfig::new() and enable only the helpers they want, or use MdxConfig::disabled() for plain Markdown. An app that wants JSX-like helpers outside this subset lowers that syntax itself, before compilation.
Unsupported helper shapes fail at build time. Code fences inside <Cards> or <Files>, for example, return a lowering error instead of producing surprising Markdown.
The example below shows what the lowerer hands to the parser.
Walking through a lowering example
Authored MDX:
<Cards>
<Card title="Runtime" description="Render a registry." href="/docs/proa-docs/runtime" />
</Cards>
<Files>
<Folder name="src">
<File name="main.rs" />
</Folder>
</Files>
Lowered Markdown:
- [Runtime](/docs/proa-docs/runtime): Render a registry.
```proa-filetree
src/
main.rs
```
After lowering, the rest of the pipeline sees Markdown and special code fence languages.
Highlighting code fences
With CodeHighlight::new(), Syntect parses recognized languages during the docs build:
let markdown = proa_docs_build::MarkdownPipeline::default()
.code_highlighting(proa_docs_build::CodeHighlight::new());
With MdxConfig::normalize_fence_metadata(true) on, the compiler removes
unsupported fence metadata before Markdown parsing while preserving
filename="..." and selected-line ranges. For example, ```rust filename="src/main.rs" title="ignored" {2,4-5} becomes ```rust filename="src/main.rs" {2,4-5}.
Filenames
Add filename="..." after the language to render a header above the code:
```rust filename="src/components/hero.rs"
pub struct Hero;
```
The generated block carries data-filename, data-code-header,
data-code-icon, and data-code-filename hooks. Recognized fence languages
use a matching language or tool icon; aliases such as rs, ts, tsx, sh,
py, yml, and mdx resolve to their canonical icon. Unknown and unlabeled
fences use a generic file icon. The filename is escaped in both attribute and
text positions, and quoted filenames may contain spaces. The existing copy
button moves into the header and still copies only the contents of
<pre><code>.
Use comma-separated line numbers and inclusive ranges to draw attention to the important lines in a block:
fn main() {
let message = "Highlighted";
println!("{message}");
println!("These ranges are inclusive");
}
The compiler emits data-code-line on every rendered line and
data-highlighted-line on the selected lines. Line wrappers preserve the
block's exact text, so the copy button still copies only the original source.
Syntect emits scopes as stable semantic classes such as hl-keyword,
hl-string, and hl-comment; the site stylesheet owns their colors. The
compiler writes no theme and no inline color into the generated HTML.
Highlight-enabled blocks carry data-code-highlight="syntect". Skip that marker
in a client-side highlighter so it never processes the generated spans twice.
Unknown languages still render as escaped <pre><code class="language-...">.
The bundled Syntect syntax set does not currently include TypeScript, TSX, or
TOML, so those fences take the same safe fallback.
Generic code blocks always include a copy button. Blocks with a filename place it in the generated header:
<div data-code-block data-lang="text" data-filename="example.txt" class="proa-docs-code-block">
<div class="proa-docs-code-header" data-code-header>
<span data-code-file-icon aria-hidden="true">...</span>
<span data-code-filename>example.txt</span>
<button type="button" data-code-copy>Copy</button>
</div>
<pre><code class="language-text">...</code></pre>
</div>
The reserved proa-filetree fence renders as a structured file tree instead of a generic code block:
```proa-filetree
app/
src/
main.rs
Cargo.toml
```
Rendered output includes stable hooks:
<div class="proa-docs-filetree" data-proa-docs-filetree role="tree">
<div class="proa-docs-filetree-row" data-kind="folder" role="treeitem" aria-expanded="true">
<code>app/</code>
</div>
</div>
Diagrams use the same fence mechanism, with a client-side renderer at the end.
Rendering Mermaid diagrams
Enable Mermaid code fences with MermaidConfig::new():
let markdown = proa_docs_build::MarkdownPipeline::default()
.mermaid(proa_docs_build::MermaidConfig::new());
With Mermaid on, a mermaid code fence emits a static placeholder:
<figure data-proa-docs-mermaid>
<pre data-proa-docs-mermaid-source hidden>...</pre>
<div data-proa-docs-mermaid-output aria-label="Mermaid diagram"></div>
</figure>
The app owns the client renderer that finds those attributes and draws the diagram. With MermaidFallback::CodeBlock, a <noscript> code block stays available.
Type tables follow the same pattern, except an external command supplies the data.
Generating TypeScript type tables
MarkdownPipeline::typescript(...) configures TypeScript docgen. The parser recognizes inline directives whose tag name matches TypeScriptDocgen::component_name, which defaults to auto-type-table:
<auto-type-table path="./src/Button.tsx" name="ButtonProps" />
The request stores:
- source Markdown path
- optional TypeScript file path
- optional type name
- optional inline type
Compilation fails when type-table requests exist and you enable TypeScript docgen without a command. That keeps unresolved API placeholders from shipping by accident.
The compiler escapes unknown inline HTML instead of rendering it. The parser intercepts only the configured type-table directive.
Whatever survives as visible text also feeds the search index.
Collecting search text
The Markdown pass records searchable text from:
- headings
- paragraph text
- inline code
- fenced code blocks
- soft breaks as spaces
Pages opt out with:
---
search: false
---
The compiler writes search documents only when the DocsSource search mode is on.
The same pass collects the other index a docs site needs: its links.
Recording link references
The renderer records hrefs from Markdown links and image sources. Link validation later checks those references against the compiled page graph:
[Runtime](/docs/proa-docs/runtime)
[Current page heading](#search-text)

The current validator skips external URLs, mailto:, and tel:. Internal routes and hashes can fail the build under LinkValidation::strict().
Alongside the HTML, the compiler keeps the Markdown that produced it.
Serving static Markdown
The generated registry stores the processed Markdown source for each page. For .md files this is the original source. For .mdx files with MdxConfig enabled, this is the lowered Markdown. render_static serves it at the page's .md route, and the marketing site builds /llms-full.txt from the same stored strings.
One authored page therefore produces three useful artifacts:
| Artifact | Purpose |
|---|---|
page.article_html | Trusted generated HTML for runtime rendering. |
page.markdown | Source Markdown for .md routes and LLM ingestion. |
page.search_text in generated search JSON | Search index input. |
Keep the pass predictable, and those three artifacts stay in agreement.
Keeping the pipeline production-safe
- Keep configured MDX lowering small and deterministic.
- Fail unsupported helper syntax at build time.
- Normalize code fence metadata before Markdown parsing when authored fences carry titles or other UI metadata.
- Prefer fenced code blocks over inline HTML examples.
- Keep
proa-filetree, Mermaid, and type-table directives documented in the authoring guide. - Run link validation after Markdown rendering so it checks heading IDs and route paths against compiler output.
Next steps
- Config reference
- Every builder on
MarkdownPipeline,MdxConfig, andMermaidConfig.
- Every builder on
- Authoring
- Write pages with the helpers this pipeline lowers.
- Link validation
- Turn recorded link references into a build gate.
- Search
- Turn the recorded search text into a working search UI.