Docs

Markdown pipeline

Follow a Markdown page through the Proa Docs build pass.

Open Markdown

proa_docs_build compiles Markdown into escaped static HTML at build time, so no parser runs when a reader loads a page. This page follows that pass through MDX lowering, code fences, Mermaid, and search text.

How does the pipeline run?

One pass records heading IDs, table-of-contents data, links, search text, code blocks, file trees, Mermaid shells, optional configured MDX helper lowering, and optional type-table requests.

pipeline.txt
source .md or configured .mdx
  -> optional MdxConfig lowering
  -> pulldown-cmark parser
  -> escaped HTML fragments
  -> TOC entries from headings
  -> link references for validation
  -> search text from visible content and code
  -> generated DocPage fields

This keeps runtime docs rendering static. The generated registry stores trusted HTML that the compiler produced, not user-provided HTML strings.

Each stage starts from what the parser recognizes.

What does the default parser enable?

The default parser enables:

The compiler escapes inline HTML unless it recognizes a configured type-table directive. It also escapes link and image URLs, and it strips javascript: and vbscript: URLs before writing HTML.

Headings get one more treatment: a stable id.

Generating heading IDs and TOC entries

Every Markdown heading gets a deterministic id:

src/docs/core/html-rendering.mdx
## Render Traits
## Render Traits

The compiler emits:

$OUT_DIR/proa_docs/docs/html-rendering.html
<h2 id="render-traits">Render Traits</h2>
<h2 id="render-traits-2">Render Traits</h2>

It also records matching TocItem values:

proa_docs/src/lib.rs
pub struct TocItem {
    pub depth: u8,
    pub id: &'static str,
    pub title: &'static str,
}

Those IDs power DocsToc, heading anchors, and link validation for #hash targets.

Authors rarely write raw Markdown alone, so the pass can lower a helper subset first.

Lowering configured MDX helpers

proa_docs_build does not execute arbitrary MDX. It can lower a small supported subset before Markdown parsing:

build.rs
let markdown = proa_docs_build::MarkdownPipeline::default()
    .mdx(proa_docs_build::MdxConfig::proa());

MdxConfig::proa() enables the helper subset this site uses:

Authored helperLowered output
<Cards> with <Card /> childrenMarkdown bullet links.
<Files> with <Folder> and <File> childrenA proa-filetree code fence that renders as a structured file tree.
<AgentTerminal />A static text code fence.
<img ...>Markdown image syntax with normalized static paths.
Layout <div> wrappersThe lowerer drops them.

That keeps the runtime registry static. Downstream apps can start from MdxConfig::new() and enable only the helpers they want, or use MdxConfig::disabled() for plain Markdown. An app that wants JSX-like helpers outside this subset lowers that syntax itself, before compilation.

Unsupported helper shapes fail at build time. Code fences inside <Cards> or <Files>, for example, return a lowering error instead of producing surprising Markdown.

The example below shows what the lowerer hands to the parser.

Walking through a lowering example

Authored MDX:

src/docs/proa-docs/example.mdx
<Cards>
  <Card title="Runtime" description="Render a registry." href="/docs/proa-docs/runtime" />
</Cards>

<Files>
  <Folder name="src">
    <File name="main.rs" />
  </Folder>
</Files>

Lowered Markdown:

$OUT_DIR/proa_docs/docs/markdown/example.md
- [Runtime](/docs/proa-docs/runtime): Render a registry.

```proa-filetree
src/
  main.rs
```

After lowering, the rest of the pipeline sees Markdown and special code fence languages.

Highlighting code fences

With CodeHighlight::new(), Syntect parses recognized languages during the docs build:

build.rs
let markdown = proa_docs_build::MarkdownPipeline::default()
    .code_highlighting(proa_docs_build::CodeHighlight::new());

With MdxConfig::normalize_fence_metadata(true) on, the compiler removes unsupported fence metadata before Markdown parsing while preserving filename="..." and selected-line ranges. For example, ```rust filename="src/main.rs" title="ignored" {2,4-5} becomes ```rust filename="src/main.rs" {2,4-5}.

Filenames

Add filename="..." after the language to render a header above the code:

src/docs/example.mdx
```rust filename="src/components/hero.rs"
pub struct Hero;
```

The generated block carries data-filename, data-code-header, data-code-icon, and data-code-filename hooks. Recognized fence languages use a matching language or tool icon; aliases such as rs, ts, tsx, sh, py, yml, and mdx resolve to their canonical icon. Unknown and unlabeled fences use a generic file icon. The filename is escaped in both attribute and text positions, and quoted filenames may contain spaces. The existing copy button moves into the header and still copies only the contents of <pre><code>.

Use comma-separated line numbers and inclusive ranges to draw attention to the important lines in a block:

fn main() {
    let message = "Highlighted";

    println!("{message}");
    println!("These ranges are inclusive");
}

The compiler emits data-code-line on every rendered line and data-highlighted-line on the selected lines. Line wrappers preserve the block's exact text, so the copy button still copies only the original source.

Syntect emits scopes as stable semantic classes such as hl-keyword, hl-string, and hl-comment; the site stylesheet owns their colors. The compiler writes no theme and no inline color into the generated HTML. Highlight-enabled blocks carry data-code-highlight="syntect". Skip that marker in a client-side highlighter so it never processes the generated spans twice.

Unknown languages still render as escaped <pre><code class="language-...">. The bundled Syntect syntax set does not currently include TypeScript, TSX, or TOML, so those fences take the same safe fallback.

Generic code blocks always include a copy button. Blocks with a filename place it in the generated header:

$OUT_DIR/proa_docs/docs/example.html
<div data-code-block data-lang="text" data-filename="example.txt" class="proa-docs-code-block">
  <div class="proa-docs-code-header" data-code-header>
    <span data-code-file-icon aria-hidden="true">...</span>
    <span data-code-filename>example.txt</span>
    <button type="button" data-code-copy>Copy</button>
  </div>
  <pre><code class="language-text">...</code></pre>
</div>

The reserved proa-filetree fence renders as a structured file tree instead of a generic code block:

src/docs/proa-docs/example.mdx
```proa-filetree
app/
  src/
    main.rs
  Cargo.toml
```

Rendered output includes stable hooks:

$OUT_DIR/proa_docs/docs/example.html
<div class="proa-docs-filetree" data-proa-docs-filetree role="tree">
  <div class="proa-docs-filetree-row" data-kind="folder" role="treeitem" aria-expanded="true">
    <code>app/</code>
  </div>
</div>

Diagrams use the same fence mechanism, with a client-side renderer at the end.

Rendering Mermaid diagrams

Enable Mermaid code fences with MermaidConfig::new():

build.rs
let markdown = proa_docs_build::MarkdownPipeline::default()
    .mermaid(proa_docs_build::MermaidConfig::new());

With Mermaid on, a mermaid code fence emits a static placeholder:

$OUT_DIR/proa_docs/docs/example.html
<figure data-proa-docs-mermaid>
  <pre data-proa-docs-mermaid-source hidden>...</pre>
  <div data-proa-docs-mermaid-output aria-label="Mermaid diagram"></div>
</figure>

The app owns the client renderer that finds those attributes and draws the diagram. With MermaidFallback::CodeBlock, a <noscript> code block stays available.

Type tables follow the same pattern, except an external command supplies the data.

Generating TypeScript type tables

MarkdownPipeline::typescript(...) configures TypeScript docgen. The parser recognizes inline directives whose tag name matches TypeScriptDocgen::component_name, which defaults to auto-type-table:

src/docs/components/button.mdx
<auto-type-table path="./src/Button.tsx" name="ButtonProps" />

The request stores:

Compilation fails when type-table requests exist and you enable TypeScript docgen without a command. That keeps unresolved API placeholders from shipping by accident.

The compiler escapes unknown inline HTML instead of rendering it. The parser intercepts only the configured type-table directive.

Whatever survives as visible text also feeds the search index.

Collecting search text

The Markdown pass records searchable text from:

Pages opt out with:

src/docs/proa-docs/example.mdx
---
search: false
---

The compiler writes search documents only when the DocsSource search mode is on.

The same pass collects the other index a docs site needs: its links.

The renderer records hrefs from Markdown links and image sources. Link validation later checks those references against the compiled page graph:

src/docs/proa-docs/example.mdx
[Runtime](/docs/proa-docs/runtime)
[Current page heading](#search-text)
![Screenshot](/static/images/docs/shell.png)

The current validator skips external URLs, mailto:, and tel:. Internal routes and hashes can fail the build under LinkValidation::strict().

Alongside the HTML, the compiler keeps the Markdown that produced it.

Serving static Markdown

The generated registry stores the processed Markdown source for each page. For .md files this is the original source. For .mdx files with MdxConfig enabled, this is the lowered Markdown. render_static serves it at the page's .md route, and the marketing site builds /llms-full.txt from the same stored strings.

One authored page therefore produces three useful artifacts:

ArtifactPurpose
page.article_htmlTrusted generated HTML for runtime rendering.
page.markdownSource Markdown for .md routes and LLM ingestion.
page.search_text in generated search JSONSearch index input.

Keep the pass predictable, and those three artifacts stay in agreement.

Keeping the pipeline production-safe

Next steps

Search

Type at least 2 characters