Compare commits

..

5 commits

Author SHA1 Message Date
randogoth
603f29b0f0 feat: gate a run of blocks to chosen output formats 2026-10-09 16:40:21 +03:00
randogoth
bb7a5111a1 docs: condense the README, fold the audit summary in, and tighten the declaration 2026-10-06 11:04:09 +03:00
randogoth
19e7cfaca3 docs: drop the predecessor comparisons from comments and tests 2026-10-06 11:04:09 +03:00
randogoth
36faa4fb93 build: package the server as an OCI container image 2026-10-06 11:04:06 +03:00
randogoth
088e860cad fix: a hard break between links is still only links
The gemtext renderer prints a bare => line for a paragraph holding nothing but
links, which is what lets a generator emit a navigable menu. is_only_links
tolerated soft breaks but not hard ones, so a generator had to choose: soft
breaks and a clean gemtext menu, or hard breaks and a menu that does not run
together on the formats that lay a paragraph out inline.

That was a false choice. A hard break between links is still a run of nothing
but links, and the inline formats already turn it into <br/>. Without this,
asking for the break made gemtext read the paragraph as prose and print its
flattened text alongside the => lines, listing every entry twice.

Found serving a real capsule: six area links rendered as one run-on line in a
browser and on a handset while gemtext looked perfect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-06 09:41:26 +03:00
30 changed files with 657 additions and 383 deletions

6
.dockerignore Normal file
View file

@ -0,0 +1,6 @@
/target
/result
/result-*
.git
.jj
.devbox

View file

@ -7,4 +7,8 @@ This format is based on [AI-DECLARATION.md](https://ai-declaration.md/en/0.1.2).
## Notes ## Notes
- - **The idea was mine, and so were the calls.** itsybitsy was conceived as a Rust server around *md2txt* and *wapdown*, my own Python renderers, with three requirements fixed before any code existed: one process serving several domains, output formats as plugins, configuration in TOML per directory. Porting the renderers rather than shelling out to them, one crate per format behind a cargo feature, flat configuration with no inheritance, which protocols shipped and in what order, TLS terminated in front of Gemini: the model proposed and argued for each; I decided.
- **The model wrote everything.** *Claude* did the Rust, the test suites and the documentation, in milestones that each ended in something runnable and verified. Parity with the originals came from diffing real output against *md2txt*, *pyfiglet* and *wapdown*. Costs of cache memory and a TLS stack were measured before implementing these features.
- **The audit found real bugs.** Three defects, the worst of them a document that could kill the whole process and every virtual host with it. All three are fixed and pinned by regression tests. See the security audit section of the README for more info.

117
AUDIT.md
View file

@ -1,117 +0,0 @@
# Security and robustness audit — 2026-10-04
Audit of itsybitsy at commit `c29364d8` ("serve gemini behind a tls terminator"), covering integration behaviour, adversarial input and load. Three defects were found and fixed in `49a63503`; the remaining items are recorded below rather than fixed, because each is a design decision rather than a bug.
The headline result is the first finding: a single content file could abort the whole process, taking every virtual host with it, in a way the existing panic handler could not catch.
## Method
A release binary built with `--features "wml figlet hyphenation"` was run against a purpose-built content tree with two sites and seven listeners — HTTP (three: negotiating, single-format, and one with `max_connections = 5`), Spartan, Gemini, Nex and Gopher. Requests were sent as raw bytes over real sockets, so each protocol's own parser was exercised rather than a client library's idea of it.
The tree deliberately contained things a well-behaved tree would not: a canary file outside both roots, symlinks pointing to that file, to `/etc/passwd` and to `/tmp`, a hidden page, files whose names carry CR LF and tab bytes, documents nested thousands of levels deep, an 8 MiB single line, a 200 000-row table, a 2.5 GB Markdown file, include cycles, a 40-deep include chain, and a 2²⁴ diamond include fan-out.
138 checks ran in four phases:
| Phase | Checks | Covers |
| --- | --- | --- |
| Integration | 59 | All five protocols end to end, negotiation, virtual hosting, every output format, card sub-documents, live reload |
| Containment | 16 | 27 traversal encodings × 5 protocols, dotfiles, symlink escapes, cross-site isolation |
| Protocol abuse | 23 | Response splitting, request smuggling, header flooding, `Host` abuse, request caps, upload abuse, slow clients |
| Exhaustion | 25 | Nesting depth, oversized documents, include bombs, include and art containment |
| Load | 15 | Concurrency, throughput, the connection cap, memory growth |
Everything of lasting value is now in the repository's own suite as 13 regression tests. The ad-hoc harness was not kept.
## Findings
### 1. Deep nesting aborted the process — critical, fixed
A served document containing 10 000 nested block quotes killed the server outright:
```
thread '<unknown>' has overflowed its stack
fatal runtime error: stack overflow, aborting
[exited with code 134]
```
Severity comes from three things together. A stack overflow raises SIGABRT rather than unwinding, so the `catch_unwind` at the handler boundary — which exists precisely so one bad document cannot take the process down — could not contain it. One process serves every configured site, so the failure is not scoped to the site whose content caused it. And nothing restarts the process on its own.
The Markdown parser was not at fault: pulldown-cmark handled 100 000 levels without trouble, being iterative. The recursion was ours, in three separate walks over the parsed document — the directive pass, each renderer's block walk, and the derived `Drop` for `Block`. The depth at which it died therefore depended on available stack, not on any limit in the code: a debug test thread died at 1 000 levels where the release server survived to 10 000.
**Fix.** A single cap at the parse boundary (`MAX_NESTING = 100` in `core/src/parse.rs`), checked where the nesting depth is already explicit as the builder's open-container stacks. Rewriting three recursive walks would have been the alternative; capping once means none of them can see a tree deep enough to matter, including any walk added later. An over-deep document is refused whole — HTTP 500, Gemini 40, Spartan 5 — rather than truncated at the limit, which would serve a page missing most of its content with nothing to indicate it.
The limit is two orders of magnitude above real prose. Ten levels of nesting is already unusual.
A secondary result worth recording, because it stops someone adding a limit that is not needed: inline nesting cannot run away. 150 consecutive `*` produce 75 levels of emphasis and 150 consecutive `[` produce one, because pulldown-cmark pairs delimiters and forbids links from nesting at all. Only block containers are unbounded. A test pins this.
### 2. A filename could inject a response header — moderate, fixed
```
GET /ev%0d%0aX-Injected:%20yes.md
HTTP/1.1 301 Moved Permanently
Location: /ev
X-Injected: yes <- injected
Connection: close
```
`url_for` interpolated the resolved filename into a URL with no encoding, and that URL is emitted as a redirect target. Spartan (`3 /ev\r\nX-Injected: yes`) and Gemini (`31 ...`) terminate their status lines the same way and split identically — one defect, three protocols.
Exploiting it requires a file named `ev\r\nX-Injected: yes.md`, which is legal on Linux. Under the current threat model the content tree is author-controlled, which is what keeps this moderate rather than critical; it becomes remotely reachable the moment any part of a tree accepts contributions from someone who is not the operator.
**Fix.** `url_for` now percent-encodes everything outside RFC 3986's unreserved set, leaving `/` as the separator. `clean_path` already decodes on the way in, so URLs round-trip, and a test asserts that. This also repaired a quieter bug that had not been noticed: filenames containing a space, `?`, `#`, `%` or non-ASCII characters previously produced URLs that were wrong or truncated. `/my%20notes`, `/a%23b` and `/caf%C3%A9` now resolve.
### 3. Large files were read before being rejected — low, fixed
`preprocess::expand` called `fs::read_to_string` on each file, allocating it in full, and only then applied the 8 MiB expansion cap. A 2.5 GB Markdown file in the tree therefore cost 2.5 GB of transient allocation per request, despite no more than 8 MiB of it ever being usable. With the default `max_connections = 256`, concurrent requests multiply that.
**Fix.** The file's length is checked before it is opened, and the read itself goes through a capped reader so the bound holds even if the file grew since the check.
## Results after the fixes
Containment held everywhere. 27 traversal encodings — double-encoded, overlong UTF-8, backslash, NUL byte, `..;/`, absolute paths — were tried against all five protocols, 135 attempts, and none returned the canary or `/etc/passwd`. Symlinks out of the root were refused on the resolved path, not the requested one. Neither site could read the other's tree, and the server configuration file, which sits outside both roots, was unreachable by every spelling tried.
No request shape broke a handler. Absolute-form and authority-form targets, missing and bogus HTTP versions, bare LF line endings, NUL bytes, invalid UTF-8, a 5 000-byte method, `Transfer-Encoding` with `Content-Length`, and a pipelined second request all produced exactly one response per connection and left the server serving. A `Host` carrying CR LF, NUL, 5 000 bytes or non-ASCII injected nothing. Every protocol enforced its request cap: Gemini 1026 bytes, Spartan 4096, Nex 2048, Gopher 512. A Spartan upload claiming 10 GB was refused in under three seconds rather than drained.
Include bombs were all bounded: the cycle, the self-reference, the 40-deep chain and the 2²⁴ diamond fan-out each produced an error and a live server, the fan-out caught by the byte cap that a per-stack cycle check cannot see.
Load behaved well:
| Measurement | Result |
| --- | --- |
| 32 concurrent clients, 3 200 requests | 8 464 req/s, 0 errors, p50 3.5 ms, p99 7.7 ms |
| 64 concurrent clients | 8 009 req/s, 0 errors |
| Mixed load across all five protocols | 9 248 req/s, 0 errors |
| 6 400 requests, leak check | RSS flat at 122 660 KiB throughout |
| 8 concurrent readers of a 4 MiB page | RSS delta 0 KiB — one shared cached copy |
| Connection cap of 5, 10 excess connections | 10/10 closed immediately; 13 threads while held, 8 when idle |
## Outstanding
**The render cache is unbounded.** Measured from a cold start: 105 pages totalling 16 588 KiB of Markdown, rendered into five formats, grew RSS from 3 876 KiB to 106 092 KiB — a cache cost of **6.2× the source size**. Memory is flat under repeated requests, so this is growth by distinct page visited rather than a leak, and at capsule scale it is fine. 6.2× is the multiplier to size `cache_max_bytes` with when LRU eviction lands.
**An oversized header produces a TCP reset instead of the 400.** The server writes the status and closes while the client is still sending, so the client may never read the response. nginx behaves similarly. Fixing it means draining a bounded amount before closing, as the Spartan listener already does for uploads.
**`ListenerSpec.width` is accepted and ignored.** A per-listener width cannot work until the render cache is keyed on `(format, width)` rather than format alone, so this is not a one-line change. Either delete the key or key the cache.
**No graceful shutdown.** SIGTERM terminates mid-response. Queued for the operations milestone along with SIGHUP reload.
## What this audit did not cover
The Gemini listener expects TLS to be terminated in front of it, and **that leg was never tested** — no `stunnel`, `ghostunnel` or `openssl` was available on the audit machine and no network to fetch one. Only the plaintext protocol behind the terminator was exercised. The `stunnel` configuration in the README is written from its documentation, not from a run.
Also out of scope: no coverage-guided fuzzing of the Markdown parser or the protocol readers, which is the right tool for the input-handling code and would likely find more than hand-written cases do. All load testing was over loopback on one machine, so the figures measure the server rather than any network. No audit of the dependency tree for known advisories. The threat model throughout assumes the content tree is author-controlled; findings 1 and 2 both become materially more serious if that stops being true, and that is the assumption most worth revisiting.
## Reproducing
The three findings are pinned by tests in the repository:
- `core/src/parse.rs` — `nesting_past_the_cap_is_refused_rather_than_overflowing_the_stack`, `deeply_repeated_inline_markers_are_bounded_by_the_parser_itself`
- `core/src/path.rs` — `url_for_escapes_bytes_that_would_end_a_response_header`, `an_encoded_url_still_resolves_back_to_the_same_path`
- `core/src/preprocess.rs` — `an_oversized_file_is_refused_without_being_read_into_memory`
- `bin/tests/listeners.rs` — `a_document_nested_past_the_cap_is_an_error_and_the_server_survives`, `an_injected_redirect_target_is_escaped_on_every_protocol`
```bash
devbox run check # 320 tests
devbox run -- cargo test --workspace --features "wml figlet hyphenation" # 375 tests
```

29
Dockerfile Normal file
View file

@ -0,0 +1,29 @@
# Multi-stage build: a slim Rust toolchain compiles the release binary, and a
# slim Debian runtime carries it with no toolchain. The default entrypoint
# serves the demo site from /srv/itsybitsy; mount a real server file and
# content roots over those paths to serve something else.
FROM rust:1.97-slim AS build
WORKDIR /src
COPY Cargo.toml Cargo.lock ./
COPY core core
COPY gemtext gemtext
COPY wap wap
COPY text text
COPY bin bin
# The demo config serves wml decks, which the `wml` feature adds to the wap
# format; the listeners the config names are validated against the registry.
RUN cargo build --release --locked --package itsybitsy --features wml
FROM debian:bookworm-slim
RUN useradd --uid 1000 --create-home itsybitsy
WORKDIR /srv/itsybitsy
COPY --from=build /src/target/release/itsybitsy /usr/local/bin/itsybitsy
# The container config binds 0.0.0.0, because a loopback bind would leave
# every published port unreachable from outside the container. --chmod keeps
# the file readable by the runtime user whatever the checkout's own modes are.
COPY --chmod=644 docker/itsybitsy.toml ./itsybitsy.toml
COPY content ./content
USER itsybitsy
EXPOSE 8080 3000 1900 1965 7070
ENTRYPOINT ["itsybitsy", "--config", "/srv/itsybitsy/itsybitsy.toml"]

236
README.md
View file

@ -1,32 +1,24 @@
# 🕸 itsybitsy # 🕸 itsybitsy
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE) [![AI-DECLARATION: copilot](https://img.shields.io/badge/%E4%B7%BC%20AI--DECLARATION-copilot-fee2e2?labelColor=fee2e2)](https://ai-declaration.md) [![License: Apache-2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE) [![AI-DECLARATION: copilot](https://img.shields.io/badge/%E4%B7%BC%20AI--DECLARATION-copilot-fee2e2?labelColor=fee2e2)](https://ai-declaration.md) ![formats](https://img.shields.io/badge/formats-gemtext%20%C2%B7%20text%20%C2%B7%20html%20%C2%B7%20XHTML--MP%20%C2%B7%20WML-4b5563) [![Built with Devbox](https://www.jetify.com/img/devbox/shield_galaxy.svg)](https://github.com/jetify-com/devbox/)
Serve folders of Markdown to different domains over HTTP, [Gemini](https://geminiprotocol.net/docs/specification.gmi), [Spartan](https://portal.mozz.us/gemini/spartan.mozz.us/), [Nex](https://nightfall.city/nex/info/specification.txt) and [Gopher](https://www.rfc-editor.org/rfc/rfc1436), rendered on request with no build step.
A Rust re-implementation of [smolweb](https://code.randogoth.com/randogoth/smolweb), which serves one folder from one process with its output formats hardcoded and its configuration in per-file frontmatter. itsybitsy changes all three: one process serves many domains, each output format is a separate crate behind a cargo feature, and configuration lives in a TOML file per content directory. Serve folders of Markdown to different domains over HTTP, [Gemini](https://geminiprotocol.net/docs/specification.gmi), [Spartan](https://portal.mozz.us/gemini/spartan.mozz.us/), [Nex](https://nightfall.city/nex/info/specification.txt) and [Gopher](https://www.rfc-editor.org/rfc/rfc1436), rendered on request with no build step. One process serves many virtual hosts, each output format is a separate crate behind a cargo feature, and configuration lives in TOML files. Gemini expects TLS to be terminated in front of it; that is the one deployment requirement itsybitsy does not satisfy by itself.
## Status The output formats are Rust ports of previous projects: the standalone Markdown to esoteric text format converter [md2txt](https://code.randogoth.com/randogoth/md2txt) and a spinoff from that for WML, [wapdown](https://code.randogoth.com/randogoth/wapdown).
It serves. All five protocols work across five output formats, with virtual hosting wherever the protocol carries a hostname. Gemini expects TLS to be terminated in front of it, which is the one deployment requirement itsybitsy does not satisfy by itself. ## Deployments
| Area | State | Itsybitsy can be seen in action running the Gemini/Spartan/Nex/Gopher content for https://smol.place, the Gemini/Spartan capsules for https://randogoth.com and https://wap.randogoth.com
| --- | --- |
| Server configuration, virtual hosts, `--check` | implemented |
| Path resolution, traversal containment, live reload | implemented |
| Markdown parsing, includes, ASCII art, card breaks | implemented |
| gemtext, XHTML-MP and HTML output | implemented |
| HTTP, Spartan and Nex listeners, virtual hosting | implemented |
| Fixed-width text output (Nex, later Gopher) | implemented |
| FIGlet banners and hyphenation, behind features | implemented |
| WML decks and card sub-documents, behind a feature | implemented |
| Gopher, with dot-stuffing and the RFC 4266 root menu | implemented |
| Gemini, behind a TLS terminator | implemented |
| Built-in TLS, so no terminator is needed | not planned for now |
## Configuration ## Usage
Two files. A server file names the virtual hosts and the listeners: ```bash
itsybitsy --config /etc/itsybitsy.toml # serve
itsybitsy --config /etc/itsybitsy.toml --check # validate configuration only
```
A server file names the sites and the listeners:
```toml ```toml
version = 1 version = 1
@ -48,167 +40,36 @@ formats = ["gemtext"]
site = "smol" site = "smol"
``` ```
A page is rendered into the union of what the enabled listeners can serve, and no more: a configuration with no HTTP listener never builds any HTML. HTTP and Spartan route by the hostname the request carries, falling back to `default_site`; Nex carries none, so its listener names one site outright. HTTP and Spartan route by the hostname the request carries, falling back to `default_site`; Nex and Gopher carry none, so their listener names one site outright. Every format a listener serves must be compiled in, which `--check` reports. A `.itsybitsy.toml` in a content directory sets rendering options for that directory — `[defaults]` for all its Markdown files, `[page."name.md"]` for one — with no inheritance from parent directories. Editing it re-renders on the next request.
A `.itsybitsy.toml` in any content directory says how the Markdown in **that directory** renders. There is no inheritance: a subdirectory without its own file uses the built-in defaults rather than its parent's. Repetition in a deep tree is the price of never having to look elsewhere to know how a page renders. ### Formats and protocols
```toml | Format | Served as | Behind feature |
[defaults] | --- | --- | --- |
max_age = 3600 | `gemtext` | `text/gemini` | default |
margin_left = 2 | `xhtmlmp`, `html` | WAP and desktop markup | default |
h1_style = "underline:=" | `text` | `text/plain`, wrapped to 80 columns | default |
| `wml` | WML 1.3 decks for WAP 1.x handsets | `wml` |
| FIGlet banners, hyphenation for `text` | | `figlet`, `hyphenation` |
[page."index.md"] HTTP negotiates formats from `Accept` and `?format=<id>` overrides it. Two decorative text features are off by default to save binary size: `cargo build --release --features "figlet hyphenation"`.
title = "Notes"
max_age = 300
```
`[defaults]` applies to every Markdown file in the directory; a `[page."name.md"]` table overrides it for one file. `title` is per-page by nature, so setting it in `[defaults]` is an error; left unset, it is derived from the page's first level-1 heading. ### Markdown
The fixed-width text format reads `margin_left`, `margin_right`, `paragraph_spacing`, `h1_style` through `h6_style`, `blockquote_bars`, `list_indent`, `code_block_line_numbers` and `wrap_code_blocks`. A heading style is `underline`, `underline:<char>`, `figlet`, `figlet:<font>`, `markers` or `plain`; by default the first three levels are underlined with `=`, `-` and `~`, and deeper ones keep their `#` markers, since underline characters run out before heading levels do. Editing the file re-renders that directory's pages on the next request, and a file that fails to parse makes them error rather than silently falling back to defaults. Markdown is parsed once per document and rendered to every format. Everything itsybitsy adds to CommonMark stays portable:
Because the name begins with a dot, the same rule that keeps `.secret.md` out of the URL space keeps this file out of it too.
`--check` validates the configuration and reports the routing it resolved, without binding a port:
```console
$ itsybitsy --config /etc/itsybitsy.toml --check
sites:
smol: /srv/smol/content
hosts:
smol.place -> smol
www.smol.place -> smol
listeners:
web: http on 0.0.0.0:8080 [xhtmlmp, html] by host, default smol
formats rendered per page: html, xhtmlmp
```
Everything the schema cannot express is checked here rather than on first request: a format no enabled feature provides, two sites claiming one hostname, a content root that does not exist, a server file sitting inside a content root where it would be served.
## Markdown
Markdown is parsed once per document into one representation that every output format reads, rather than once per format. Everything itsybitsy adds to CommonMark is either a construct other tools already understand or invisible to them, so a document stays portable:
| Written | Means | | Written | Means |
| --- | --- | | --- | --- |
| `![[path]]` | splice that file's lines in here | | `![[path]]` | splice that file's lines in here |
| `![alt](art.txt)` | an image whose target is a text file, inlined verbatim | | `![alt](art.txt)` | an image whose target is a text file, inlined verbatim |
| `<!-- card Title -->` | a card divider, for formats that paginate | | `<!-- card Title -->` | a card divider, for formats that paginate |
| `---` | an untitled divider |
| `<!-- center -->`, `<!-- right margin=4 -->` | alignment for the block that follows | | `<!-- center -->`, `<!-- right margin=4 -->` | alignment for the block that follows |
| `<!-- only gemtext text -->` … `<!-- end -->` | a run of blocks only those formats get |
| `<!-- except wml -->` … `<!-- end -->` | a run of blocks every other format gets |
Only the include is non-standard, and it is the spelling Obsidian and its relatives established; CommonMark renders it as literal text, which is why it is the one construct handled before parsing. The rest are ordinary HTML comments and images, interpreted after parsing — so an HTML comment that is not a recognised directive stays a comment, at the cost of a misspelled one doing nothing rather than complaining. A gate names format ids, not protocols: `gemtext` is what Gemini, Spartan, Nex and Gopher all serve, so there is no "only on Gemini". Gates nest, each one closes inside the quote or list item it was opened in, and one left unclosed is refused rather than run on to the end of the page. An id this build does not provide matches nothing, so `only wml` is hidden everywhere when the `wml` feature is off — and so is a misspelled id.
Alignment only reaches the fixed-width text formats; gemtext and HTML have no way to express it. An art label may carry its own `:center`, `:left` or `:right` token, which is removed from the label so it does not leak into the fallback text a gemtext or HTML client shows. Links should be root-relative and extensionless (`[about](/about)`), so the same link resolves identically from every protocol:
Include and art targets resolve relative to the file holding the reference and must stay inside the content root. Expansion is bounded on three axes — nesting depth, total lines, and total bytes — because a cycle check alone does not stop a long chain, and neither stops a diamond where two branches include the same file without ever repeating one on a single path.
### Differences from smolweb
Configuration keys are not carried over verbatim either: md2txt accepts three spellings of `paragraph_spacing`, and its `cache_control` is named for the header it lands in rather than the value it holds, which here is `max_age`. Its `text` and `nex` renderers differ only in heading decoration and in whether links are inlined or numbered — and the numbered form never writes the reference list its numbers point at, so there is one text renderer here rather than two. Tables, which both of its text renderers drop entirely, are rendered as aligned columns.
wapdown's own parser produces neither tables nor lists, so its WML decks render a bullet list as literal text and drop tables entirely; both come out as real markup here, lists as marked lines and tables as `<table>`, which WML has.
Gemini diverges in three places. smolweb answers 59 (bad request) to any URL that is not `gemini://`, which tells a client its request was malformed when an `https://` URL is merely for somewhere this server does not fetch from; that is 53 (proxy request refused) here, and 59 is kept for a URL that genuinely does not parse. smolweb builds absolute redirect targets from the port it is bound to, which is wrong behind a terminator; these are relative. And smolweb terminates TLS itself with a certificate named per site, so `tls` is not a configuration key here at all — naming one is an error rather than a setting that quietly does nothing.
smolweb inherits five invented directives from its two renderer libraries: `{.card}`, `{.include}`, `![[ ]]`, `#[label](art.txt)` and MultiMarkdown attribute lists (`{: .center}`) — two incompatible brace grammars, with art alignment expressible two different ways in one line. None of it appeared in real content, so itsybitsy expresses the same capabilities with standard constructs instead. Parsing once rather than once per library also means the text formats gain setext headings and the WAP formats gain tables and ASCII art, none of which their original parser handled.
## Output formats
Each format is a crate implementing one trait, compiled in behind a cargo feature. A renderer is handed the parsed document and may not open files or sockets, which is what keeps every path check in one place rather than in each format.
| Format | Served as | Notes |
| --- | --- | --- |
| `gemtext` | `text/gemini` | Unwrapped; Gemini and Spartan clients wrap for themselves |
| `xhtmlmp` | `application/vnd.wap.xhtml+xml` | Well-formed XML with the Mobile Profile doctype |
| `html` | `text/html` | The same markup without the XML declaration |
| `text` | `text/plain` | Wrapped to 80 columns, because Nex and Gopher clients do not wrap |
| `wml` | `text/vnd.wap.wml` | WAP 1.x decks, paginated to a per-card byte budget (feature `wml`) |
HTTP negotiates between them from `Accept`. Only a literal media type counts as a match, so a browser's `*/*` can never be read as willingness to receive a WAP format; `?format=<id>` overrides negotiation outright, and naming an id the listener does not serve is a bad request rather than a silent fallback.
A third party adds a format by writing a crate that depends on `itsybitsy-core`, adding it to the workspace and one `#[cfg]` arm in the binary's registry, then naming its id in `formats`. Every listener's formats are checked against the registry at startup, so a format switched off by a feature is a configuration error naming the missing id rather than a server error on the first request.
## Optional features
Two decorative capabilities for the text format are off by default, because most pages leave them alone and both cost binary size:
| Feature | Adds | Cost |
| --- | --- | --- |
| `wml` | WML 1.3 decks for WAP 1.x handsets | +37 KiB |
| `figlet` | `h1_style = "figlet"` banners, fonts `small` and `standard` | +60 KiB |
| `hyphenation` | `hyphenate = true`, English patterns | +127 KiB |
| `hyphenation-all` | every language the pattern crate carries | +3.0 MiB |
```bash
cargo build --release --features "figlet hyphenation"
```
A banner that will not fit the line, names a font this build lacks, or is asked for by a build without `figlet` falls back to the level's underline — a heading that cannot be decorated should not be lost. An unknown font name is logged, since that is almost always a typo. Likewise `hyphenate` has no effect without the feature, and an unknown `hyphen_lang` leaves the text unhyphenated; a ragged right edge is the plain-text convention anyway.
## Protocols
| Protocol | Routes by | Notes |
| --- | --- | --- |
| HTTP | `Host` header | `GET` and `HEAD`, negotiated by `Accept` |
| Spartan | the host field of its request line | One format per listener |
| Nex | nothing; its listener names one site | No status line at all |
| Gemini | the authority of the URL it is sent | TLS terminated in front; one format per listener |
| Gopher | nothing; its listener names one site | Item type 0 and the root menu |
A protocol with no hostname in its requests cannot be routed by one, so its listener names a single site outright and configuration refuses to leave that implicit once more than one site exists. HTTP, Gemini and Spartan fall back to `default_site` when a request names a host this server does not know.
Neither Nex nor Gopher has a redirect status, so a canonical target is resolved server-side and the real content comes back on the first request rather than a bounce.
### Gemini and TLS
itsybitsy does not speak TLS. The Gemini listener expects plaintext, so a TLS wrapper terminates in front of it and forwards to a loopback port:
```toml
[listener.gemini]
protocol = "gemini"
# The terminator owns 1965; this is its back end.
bind = "127.0.0.1:11965"
formats = ["gemtext"]
default_site = "smol"
```
```ini
; /etc/stunnel/gemini.conf
[gemini]
accept = 1965
connect = 127.0.0.1:11965
cert = /var/lib/itsybitsy/gemini.pem
```
`stunnel` and `ghostunnel` are built for exactly this and support SNI, so each virtual host can present its own certificate. Use a long-lived self-signed certificate rather than an ACME one: Gemini clients pin the fingerprint they first saw, and a renewal that changes it is indistinguishable from an interception.
Two things follow from terminating in front, and neither is a limitation that can be configured away. Every connection appears to come from the terminator, so request logs show its address unless it speaks the PROXY protocol, which itsybitsy does not yet read. And a client certificate never reaches us, so statuses 60 to 62 are not implementable — itsybitsy serves static documents and has nothing to authenticate, so only the logs are poorer for it.
Redirects are relative references, which the spec permits and which are the only correct form here: the port this listener is bound to is the terminator's back end, not a port any client reached, so an absolute URL built from it would point somewhere unpublished.
Gopher serves item type 0 (text) and one menu: a prefix-less `gopher://host/` means item type 1 by RFC 4266, so an empty selector gets a one-item menu pointing at `/` rather than the root document, which would be the wrong type. Text items are dot-stuffed and terminated with a lone dot; binary items are sent raw, since dot-stuffing would corrupt them and a terminator would become part of the file. There are no generated directory listings and no type 7 search.
## Card sub-documents
WML is the one output that paginates: a WAP 1.x handset has a hard per-card byte budget and refuses a deck that exceeds it. So a long page becomes a chain of screens, and a page with dividers becomes a menu card linking to the rest.
With `deck_per_card`, each card also gets its own URL under the page's:
| Request | WML client | Any other format |
| --- | --- | --- |
| `/trail` | the menu deck | the whole page |
| `/trail/weather` | that card's own deck | redirects to `/trail` |
| `/trail/nonexistent` | redirects to `/trail` | redirects to `/trail` |
| `/about/whatever` | not found | not found |
That URL space belongs to the format that claims it. A format which does not address a sub-document redirects to the parent rather than substituting something else, and a page with no dividers has no such URLs at all — `/about/whatever` is not a sub-document just because `/about` exists. Nex and Gopher have no redirect status, so they resolve to the parent's content directly instead of bouncing.
The deck keys are `split_level`, `split_on_rule`, `max_card_bytes`, `menu`, `menu_style`, `deck_per_card`, `template_nav`, `nav_next_label`, `nav_prev_label`, `nav_back_label`, `home_label` and `images`.
## URLs
Links should be root-relative and extensionless (`[about](/about)`, not `about.md`), so the same link resolves identically from every protocol.
| Request | Serves | | Request | Serves |
| --- | --- | | --- | --- |
@ -217,11 +78,46 @@ Links should be root-relative and extensionless (`[about](/about)`, not `about.m
| `/foo.md` | redirects to `/foo` | | `/foo.md` | redirects to `/foo` |
| `/img.png` | the file itself, by media type | | `/img.png` | the file itself, by media type |
Nothing outside the content root is reachable. A request target is percent-decoded before it is normalised, so an encoded `..` becomes a real one and gets clamped at the root rather than quietly matching nothing; any path component beginning with a dot is refused outright; and whatever survives is canonicalised and required to still be inside the root, which is what defeats a symlink pointing out of it. There are no generated directory listings — only `index.md`. Nothing outside a content root is reachable, dot-prefixed components are refused, and include expansion is depth-, line- and byte-capped. There are no generated directory listings — only `index.md`.
URLs the server generates are percent-encoded. A file name can contain a space, a `?`, a `#` or even a CR LF, and these URLs are emitted as redirect targets, so an unencoded one would truncate the path or end the response header and let a crafted file name inject a header of its own into every protocol. ### Gemini and TLS
A document is also refused if it nests blocks more than 100 levels deep, or if it is larger than the 8 MiB an expansion may produce. Both are about the stack and the heap rather than about the content: every pass over a parsed document recurses, and a stack overflow aborts the process instead of unwinding, so one pathological file would take every virtual host down with it. Includes are separately capped at 16 levels and 200 000 lines, with the byte cap catching the diamond fan-out that a per-stack cycle check cannot see. itsybitsy does not speak TLS. A TLS wrapper (`stunnel`, `ghostunnel`) or a load balancer owns 1965 and forwards plaintext to a loopback port the Gemini listener binds; it supports SNI, so each virtual host can present its own certificate. Use a long-lived self-signed certificate rather than an ACME one: Gemini clients pin the fingerprint they first saw. A TLS terminator in front is also how the container deployment below runs.
## Deployment
The image builds from the repository root and runs on any OCI container host:
```bash
docker build -t itsybitsy .
docker run -p 8080:8080 -p 3000:3000 -p 1900:1900 -p 7070:7070 itsybitsy
```
It serves the demo site with `docker/itsybitsy.toml`, which binds `0.0.0.0` where the checkout's own server file binds loopback, so published ports are reachable. A real site mounts over those paths:
```bash
docker run -p 8080:8080 \
-v /srv/mysite.toml:/srv/itsybitsy/itsybitsy.toml:ro \
-v /srv/mysite:/srv/content:ro itsybitsy
```
The container runs as an unprivileged user. Port 1965 stays unexposed on the host: the Gemini listener is plaintext and a TLS terminator must own it in front, as always.
The flake deploys without Docker:
```bash
nix run . -- --config ./itsybitsy.toml # build and serve
nix build .#static # fully static musl binary
scp result/bin/itsybitsy host:/usr/local/bin/
```
`packages.static` links against musl with no C dependency, so the one binary runs unchanged on any Linux host, including one that has never seen Nix; `nix run .#release-static` publishes it as a release asset.
## Security audit
An adversarial audit of integration, containment, protocol abuse and load (2026-10-04, 138 checks over raw sockets) found three defects, all fixed with regression tests: deeply nested documents overflowed the stack and aborted the process (parsing now refuses past 100 levels), a file name with CR LF could inject a response header via redirect targets (generated URLs are now percent-encoded), and a large Markdown file was read into memory before the 8 MiB expansion cap rejected it (reads are now capped). After the fixes, 135 traversal attempts across all five protocols returned nothing outside the content roots and include bombs were bounded, at roughly 8–9k req/s with flat memory.
Outstanding: the render cache is unbounded (6.2× the source size across formats), an oversized request header yields a TCP reset instead of a 400, a listener's `width` is ignored, and there is no graceful shutdown. The TLS leg in front of Gemini was not tested, and the threat model assumes the content tree is author-controlled.
## Development ## Development

View file

@ -99,7 +99,6 @@ struct Request<'a> {
/// 53 is "a resource at a domain not served by the server" and 59 is a request /// 53 is "a resource at a domain not served by the server" and 59 is a request
/// the server could not parse: a URL that parses but names another scheme or /// the server could not parse: a URL that parses but names another scheme or
/// host is a proxy request, while one that does not parse is a bad request. /// host is a proxy request, while one that does not parse is a bad request.
/// smolweb answers 59 to both, including to an ordinary `https://` URL.
fn parse(line: &str) -> Result<Request<'_>, Refusal> { fn parse(line: &str) -> Result<Request<'_>, Refusal> {
let Some((scheme, rest)) = line.split_once("://") else { let Some((scheme, rest)) = line.split_once("://") else {
return Err(Refusal { status: 59, meta: "Bad request: an absolute URL is required" }); return Err(Refusal { status: 59, meta: "Bad request: an absolute URL is required" });
@ -195,8 +194,8 @@ mod tests {
#[test] #[test]
fn another_scheme_is_a_proxy_request_not_a_parse_failure() { fn another_scheme_is_a_proxy_request_not_a_parse_failure() {
// smolweb answers 59 here, which tells a client its request was malformed // 59 would tell a client its request was malformed when it was merely for
// when it was merely for somewhere this server does not fetch from. // somewhere this server does not fetch from.
assert_eq!(refused("https://example.org/"), 53); assert_eq!(refused("https://example.org/"), 53);
assert_eq!(refused("gopher://example.org/"), 53); assert_eq!(refused("gopher://example.org/"), 53);
} }

View file

@ -16,7 +16,7 @@ use itsybitsy_core::site::{Resolution, Resource};
use crate::proto::for_log; use crate::proto::for_log;
use crate::serve::Listener; use crate::serve::Listener;
/// Gopher selectors are short; the cap is smolweb's. /// Gopher selectors are short, and nothing in the protocol needs a long one.
const MAX_REQUEST: usize = 512; const MAX_REQUEST: usize = 512;
pub fn serve(listener: &Listener, mut stream: TcpStream) -> Result<()> { pub fn serve(listener: &Listener, mut stream: TcpStream) -> Result<()> {

View file

@ -32,7 +32,8 @@ pub fn serve(listener: &Listener, mut stream: TcpStream) -> Result<()> {
if !version.starts_with("HTTP/1.") { if !version.starts_with("HTTP/1.") {
return bad_request(&mut stream); return bad_request(&mut stream);
} }
// Origin-form only. smolweb's `urlsplit` would have mishandled the others. // Origin-form only; the other request-target forms are refused, not
// half-supported.
if !target.starts_with('/') { if !target.starts_with('/') {
return bad_request(&mut stream); return bad_request(&mut stream);
} }
@ -168,8 +169,8 @@ pub fn serve(listener: &Listener, mut stream: TcpStream) -> Result<()> {
write_head(&mut stream, 200, "OK", media_type, meta.len(), &[])?; write_head(&mut stream, 200, "OK", media_type, meta.len(), &[])?;
if !head_only { if !head_only {
let mut file = std::fs::File::open(&path)?; let mut file = std::fs::File::open(&path)?;
// Streamed, not buffered: smolweb reads the whole file into // Streamed, not buffered, so a large file costs the copy buffer
// memory on every request. // rather than its own size on every request.
std::io::copy(&mut file, &mut stream)?; std::io::copy(&mut file, &mut stream)?;
} }
Ok(()) Ok(())

View file

@ -13,7 +13,7 @@ use itsybitsy_core::site::{Resolution, Resource};
use crate::proto::{for_log, read_line_capped}; use crate::proto::{for_log, read_line_capped};
use crate::serve::Listener; use crate::serve::Listener;
/// Nex requests are a single path; the cap is smolweb's. /// Nex requests are a single path, so the cap only has to be generous for one.
const MAX_REQUEST: usize = 2048; const MAX_REQUEST: usize = 2048;
pub fn serve(listener: &Listener, mut stream: TcpStream) -> Result<()> { pub fn serve(listener: &Listener, mut stream: TcpStream) -> Result<()> {

View file

@ -1,7 +1,7 @@
//! Spartan: `HOST PATH LENGTH` in, a one-digit status and a body out. //! Spartan: `HOST PATH LENGTH` in, a one-digit status and a body out.
//! //!
//! The host field is what smolweb discards; here it selects the virtual host, //! The host field selects the virtual host, falling back to the listener's
//! falling back to the listener's `default_site` when it names nothing known. //! `default_site` when it names nothing known.
use std::io::{BufReader, Read, Write}; use std::io::{BufReader, Read, Write};
use std::net::TcpStream; use std::net::TcpStream;

View file

@ -171,7 +171,7 @@ fn accept_loop(listener: Arc<Listener>, socket: TcpListener) {
}; };
// At the cap, refuse immediately rather than queueing threads without // At the cap, refuse immediately rather than queueing threads without
// bound. smolweb has no cap at all. // bound.
let open = listener.open.fetch_add(1, Ordering::SeqCst); let open = listener.open.fetch_add(1, Ordering::SeqCst);
if open >= listener.max_connections { if open >= listener.max_connections {
listener.open.fetch_sub(1, Ordering::SeqCst); listener.open.fetch_sub(1, Ordering::SeqCst);
@ -194,8 +194,9 @@ fn handle(listener: &Listener, stream: TcpStream) {
} }
// One malformed document must not take the process down, so a panic in a // One malformed document must not take the process down, so a panic in a
// handler is caught and logged. This is the Rust equivalent of smolweb's // handler is caught and logged with its cause rather than lost. A stack
// bare `except Exception`, except the cause is recorded rather than lost. // overflow is not a panic and aborts regardless, which is why the parser caps
// how deeply a document may nest.
let caught = let caught =
std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| match listener.protocol { std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| match listener.protocol {
Protocol::Http => proto::http::serve(listener, stream), Protocol::Http => proto::http::serve(listener, stream),

View file

@ -348,7 +348,7 @@ fn spartan_serves_gemtext() {
} }
#[test] #[test]
fn spartan_routes_by_the_host_field_smolweb_discards() { fn spartan_routes_by_the_host_field() {
let server = Server::start(); let server = Server::start();
assert!(server.send("spartan", b"one.test / 0\r\n").contains("# One")); assert!(server.send("spartan", b"one.test / 0\r\n").contains("# One"));
assert!(server.send("spartan", b"two.test / 0\r\n").contains("# Two")); assert!(server.send("spartan", b"two.test / 0\r\n").contains("# Two"));
@ -560,8 +560,8 @@ fn nothing_outside_the_root_is_reachable_over_any_protocol() {
// -- Gopher -------------------------------------------------------------- // -- Gopher --------------------------------------------------------------
// //
// Ported from smolweb's tests/test_gopher.py, where the exact wire bytes are the // The exact wire bytes are the assertion here: Gopher has no status line, so
// assertion: Gopher has no status line, so framing is all a client has. // framing is all a client has.
#[test] #[test]
fn an_empty_selector_gets_a_one_item_menu() { fn an_empty_selector_gets_a_one_item_menu() {
@ -591,9 +591,8 @@ fn a_text_item_ends_with_the_lone_dot_terminator() {
#[test] #[test]
fn a_missing_selector_is_a_well_formed_text_item() { fn a_missing_selector_is_a_well_formed_text_item() {
// No status to report with, so the error is the item's content. smolweb says // No status to report with, so the error is the item's content. One wording
// "Not found." here and "Not found" over HTTP and Nex; one wording is used // is used across every protocol, so this matches the HTTP and Nex bodies.
// across every protocol instead.
let server = Server::start(); let server = Server::start();
assert_eq!(server.send("gopher", b"/nope\r\n"), "Not found\n.\r\n"); assert_eq!(server.send("gopher", b"/nope\r\n"), "Not found\n.\r\n");
} }

View file

@ -3,8 +3,7 @@
//! Invalidation is by modification time and length, which is what makes editing //! Invalidation is by modification time and length, which is what makes editing
//! a file enough to see the change on the next request with no watcher and no //! a file enough to see the change on the next request with no watcher and no
//! restart. Values are built outside the lock, so two first hits on one file can //! restart. Values are built outside the lock, so two first hits on one file can
//! both build it; the work is idempotent and the second insert wins, which is //! both build it; the work is idempotent and the second insert wins.
//! the trade the Python makes too.
use std::collections::HashMap; use std::collections::HashMap;
use std::fs; use std::fs;
@ -109,9 +108,9 @@ mod tests {
use super::*; use super::*;
/// Force a modification time change. Filesystem granularity is coarse /// Force a modification time change. Filesystem granularity is coarse enough
/// enough that two writes in one test can otherwise share a timestamp, /// that two writes in one test can otherwise share a timestamp, which would
/// which is why the Python's live-reload test does the same thing. /// make a stale value look correctly cached.
fn bump_mtime(path: &Path) { fn bump_mtime(path: &Path) {
let later = SystemTime::now() + Duration::from_secs(5); let later = SystemTime::now() + Duration::from_secs(5);
File::options() File::options()
@ -137,7 +136,8 @@ mod tests {
let first = cache.get_or_insert_with(&path, Stamp::of(&path).unwrap(), build).unwrap(); let first = cache.get_or_insert_with(&path, Stamp::of(&path).unwrap(), build).unwrap();
let second = cache.get_or_insert_with(&path, Stamp::of(&path).unwrap(), build).unwrap(); let second = cache.get_or_insert_with(&path, Stamp::of(&path).unwrap(), build).unwrap();
// Identity, not just equality: the Python asserts `first is second`. // Identity, not equality: a repeat hit must return the cached value
// rather than an equal rebuild.
assert!(Arc::ptr_eq(&first, &second)); assert!(Arc::ptr_eq(&first, &second));
assert_eq!(builds.load(Ordering::Relaxed), 1); assert_eq!(builds.load(Ordering::Relaxed), 1);
} }

View file

@ -670,8 +670,8 @@ impl Default for PageSettings {
// the terminal edge. // the terminal edge.
margin_left: 2, margin_left: 2,
margin_right: 2, margin_right: 2,
// One blank line between blocks. md2txt uses two, which reads as // One blank line between blocks; two reads as double-spaced
// double-spaced throughout. // throughout.
paragraph_spacing: 1, paragraph_spacing: 1,
heading_styles: DEFAULT_HEADING_STYLES, heading_styles: DEFAULT_HEADING_STYLES,
blockquote_bars: true, blockquote_bars: true,

View file

@ -7,6 +7,8 @@
//! | --- | --- | //! | --- | --- |
//! | `<!-- card Title -->` | a card divider; invisible to any other renderer | //! | `<!-- card Title -->` | a card divider; invisible to any other renderer |
//! | `<!-- center -->` | alignment for the block that follows | //! | `<!-- center -->` | alignment for the block that follows |
//! | `<!-- only gemtext -->` … `<!-- end -->` | a run of blocks only these output formats get |
//! | `<!-- except wml -->` … `<!-- end -->` | a run of blocks every other format gets |
//! | `![alt](art.txt)` | an image whose target is text, inlined verbatim | //! | `![alt](art.txt)` | an image whose target is text, inlined verbatim |
//! //!
//! An HTML comment that is not a recognised directive stays a comment. Comments //! An HTML comment that is not a recognised directive stays a comment. Comments
@ -37,19 +39,32 @@ pub fn apply(doc: &mut Doc, base: &Path, root: &Path) -> Result<(), Error> {
fn rewrite(blocks: Vec<Block>, base: &Path, root: &Path) -> Result<Vec<Block>, Error> { fn rewrite(blocks: Vec<Block>, base: &Path, root: &Path) -> Result<Vec<Block>, Error> {
let mut out = Vec::with_capacity(blocks.len()); let mut out = Vec::with_capacity(blocks.len());
let mut pending: Option<Directive> = None; let mut pending: Option<Directive> = None;
// Gates in the order they were opened. Each recursion gets its own, so a
// gate opened inside a quote or a list item has to close inside it too.
let mut gates: Vec<Gate> = Vec::new();
for block in blocks { for block in blocks {
if let Block::Html(html) = &block { if let Block::Html(html) = &block {
match directive(html) { match directive(html) {
Some(Directive::Card { title }) => { Some(Directive::Card { title }) => {
out.push(Block::CardBreak { title }); out.extend(gated(&gates, vec![Block::CardBreak { title }]));
continue; continue;
} }
Some(align @ Directive::Align { .. }) => { Some(align @ Directive::Align { .. }) => {
pending = Some(align); pending = Some(align);
continue; continue;
} }
None => {} Some(Directive::Only(gate)) => {
gates.push(gate);
continue;
}
Some(Directive::End) if !gates.is_empty() => {
gates.pop();
continue;
}
// A close with no gate open is just a comment, as is anything
// else unrecognised.
Some(Directive::End) | None => {}
} }
} }
@ -70,24 +85,68 @@ fn rewrite(blocks: Vec<Block>, base: &Path, root: &Path) -> Result<Vec<Block>, E
other => vec![other], other => vec![other],
}; };
match pending.take() { let aligned = match pending.take() {
Some(Directive::Align { align, margin }) => out.extend(produced.into_iter().map(|b| { Some(Directive::Align { align, margin }) => produced
// A block that already carries its own alignment keeps it: the .into_iter()
// more specific marker wins over the one that precedes it. .map(|b| {
match b { // A block that already carries its own alignment keeps it:
aligned @ Block::Aligned { .. } => aligned, // the more specific marker wins over the one before it.
other => Block::Aligned { align, margin, block: Box::new(other) }, match b {
} aligned @ Block::Aligned { .. } => aligned,
})), other => Block::Aligned { align, margin, block: Box::new(other) },
_ => out.extend(produced), }
} })
.collect(),
_ => produced,
};
out.extend(gated(&gates, aligned));
}
if !gates.is_empty() {
return Err(Error::directive(
"an <!-- only --> or <!-- except --> gate is never closed; add <!-- end -->",
));
} }
Ok(out) Ok(out)
} }
/// Wrap each block in every open gate, the first opened outermost, so nested
/// gates compose: a block has to satisfy all of them to survive filtering.
fn gated(gates: &[Gate], blocks: Vec<Block>) -> Vec<Block> {
if gates.is_empty() {
return blocks;
}
blocks
.into_iter()
.map(|block| {
gates.iter().rev().fold(block, |inner, gate| Block::Gated {
formats: gate.formats.clone(),
negated: gate.negated,
block: Box::new(inner),
})
})
.collect()
}
enum Directive { enum Directive {
Card { title: Option<String> }, Card {
Align { align: Align, margin: Option<u16> }, title: Option<String>,
},
Align {
align: Align,
margin: Option<u16>,
},
/// Opens a run of blocks restricted to some output formats.
Only(Gate),
/// Closes the innermost open gate.
End,
}
/// An open gate's condition: format ids, and whether they are the formats to
/// keep (`only`) or the ones to leave out (`except`).
struct Gate {
formats: Vec<String>,
negated: bool,
} }
/// Recognise a directive in the text of an HTML block, or `None` for an ordinary /// Recognise a directive in the text of an HTML block, or `None` for an ordinary
@ -100,11 +159,31 @@ fn directive(html: &str) -> Option<Directive> {
let title = words.collect::<Vec<_>>().join(" "); let title = words.collect::<Vec<_>>().join(" ");
return Some(Directive::Card { title: (!title.is_empty()).then_some(title) }); return Some(Directive::Card { title: (!title.is_empty()).then_some(title) });
} }
if let Some(negated) = gate_named(name) {
// A gate naming no format would hide or reveal everything by accident,
// so an argument-less one stays a comment.
let formats = words.map(str::to_string).collect::<Vec<_>>();
return (!formats.is_empty()).then_some(Directive::Only(Gate { formats, negated }));
}
if name == "end" {
// `<!-- end of the list -->` is prose, not a close.
return words.next().is_none().then_some(Directive::End);
}
let align = align_named(name)?; let align = align_named(name)?;
let margin = words.find_map(|word| word.strip_prefix("margin=")?.parse().ok()); let margin = words.find_map(|word| word.strip_prefix("margin=")?.parse().ok());
Some(Directive::Align { align, margin }) Some(Directive::Align { align, margin })
} }
/// Whether a name opens a gate, and if so whether it names the formats to leave
/// out rather than the ones to keep.
fn gate_named(name: &str) -> Option<bool> {
match name {
"only" => Some(false),
"except" => Some(true),
_ => None,
}
}
fn align_named(name: &str) -> Option<Align> { fn align_named(name: &str) -> Option<Align> {
match name { match name {
"left" => Some(Align::Left), "left" => Some(Align::Left),
@ -307,6 +386,128 @@ mod tests {
assert_eq!(doc.blocks.len(), 1); assert_eq!(doc.blocks.len(), 1);
} }
// -- Format gates ------------------------------------------------------
fn gate(formats: &[&str], negated: bool, block: Block) -> Block {
Block::Gated {
formats: formats.iter().map(|id| id.to_string()).collect(),
negated,
block: Box::new(block),
}
}
#[test]
fn a_gate_wraps_every_block_of_its_run() {
let tree = Tree::new();
let doc = tree.doc("<!-- only html -->\n\nA.\n\nB.\n\n<!-- end -->\n\nC.\n").unwrap();
assert_eq!(
doc.blocks,
vec![
gate(&["html"], false, para("A.")),
gate(&["html"], false, para("B.")),
// Past the close the run is over.
para("C."),
]
);
}
#[test]
fn an_except_gate_names_the_formats_to_leave_out() {
let tree = Tree::new();
let doc = tree.doc("<!-- except wml text -->\nA.\n<!-- end -->\n").unwrap();
assert_eq!(doc.blocks, vec![gate(&["wml", "text"], true, para("A."))]);
}
#[test]
fn nested_gates_compose_with_the_first_opened_outermost() {
let tree = Tree::new();
let doc = tree
.doc("<!-- only html xhtmlmp -->\n<!-- except xhtmlmp -->\nA.\n<!-- end -->\n<!-- end -->\n")
.unwrap();
assert_eq!(
doc.blocks,
vec![gate(&["html", "xhtmlmp"], false, gate(&["xhtmlmp"], true, para("A.")))]
);
}
#[test]
fn a_card_divider_inside_a_gate_is_gated_too() {
// Otherwise WML would paginate at a divider meant for another format.
let tree = Tree::new();
let doc = tree.doc("<!-- only wml -->\n<!-- card Next -->\n<!-- end -->\n").unwrap();
assert_eq!(
doc.blocks,
vec![gate(&["wml"], false, Block::CardBreak { title: Some("Next".into()) })]
);
}
#[test]
fn an_alignment_before_a_gate_ends_up_inside_it() {
let tree = Tree::new();
let doc = tree.doc("<!-- center -->\n<!-- only text -->\nA.\n<!-- end -->\n").unwrap();
assert_eq!(
doc.blocks,
vec![gate(
&["text"],
false,
Block::Aligned { align: Align::Center, margin: None, block: Box::new(para("A.")) },
)]
);
}
#[test]
fn a_gate_may_be_opened_inside_a_quote_or_a_list_item() {
let tree = Tree::new();
let doc = tree
.doc("> <!-- only text -->\n> A.\n> <!-- end -->\n\n- <!-- only text -->\n B.\n <!-- end -->\n")
.unwrap();
let Block::BlockQuote(inner) = &doc.blocks[0] else { panic!("expected a quote") };
assert_eq!(inner, &vec![gate(&["text"], false, para("A."))]);
let Block::List { items, .. } = &doc.blocks[1] else { panic!("expected a list") };
assert_eq!(items[0], vec![gate(&["text"], false, para("B."))]);
}
#[test]
fn a_gate_left_open_is_refused() {
// Running it silently to the end of the document would hide the rest of
// the page from some formats with nothing to notice it by.
let tree = Tree::new();
let err = tree.doc("<!-- only html -->\nA.\n").unwrap_err();
assert!(err.to_string().contains("never closed"), "{err}");
}
#[test]
fn a_gate_must_close_inside_the_quote_it_was_opened_in() {
let tree = Tree::new();
assert!(tree.doc("> <!-- only html -->\n> A.\n\n<!-- end -->\n").is_err());
}
#[test]
fn a_close_with_nothing_open_stays_a_comment() {
let tree = Tree::new();
let doc = tree.doc("<!-- end -->\nA.\n").unwrap();
assert!(matches!(&doc.blocks[0], Block::Html(html) if html.contains("end")));
assert_eq!(doc.blocks[1], para("A."));
}
#[test]
fn a_gate_naming_no_format_stays_a_comment() {
// It would otherwise hide or reveal everything by accident.
let tree = Tree::new();
let doc = tree.doc("<!-- only -->\nA.\n").unwrap();
assert!(matches!(&doc.blocks[0], Block::Html(_)));
assert_eq!(doc.blocks[1], para("A."));
}
#[test]
fn a_close_carrying_prose_stays_a_comment() {
let tree = Tree::new();
let doc =
tree.doc("<!-- only html -->\nA.\n<!-- end of the gate -->\n<!-- end -->\n").unwrap();
let Block::Gated { block, .. } = &doc.blocks[1] else { panic!("expected a gated block") };
assert!(matches!(block.as_ref(), Block::Html(html) if html.contains("end of the gate")));
}
// -- Art --------------------------------------------------------------- // -- Art ---------------------------------------------------------------
#[test] #[test]

View file

@ -31,6 +31,12 @@ pub enum Error {
/// stack overflow aborts the process rather than unwinding. /// stack overflow aborts the process rather than unwinding.
TooDeep { limit: usize }, TooDeep { limit: usize },
/// A directive the parser could read but not complete, such as an
/// `<!-- only ... -->` gate that is never closed. Refused rather than
/// guessed, because the guess would hide content from some formats and show
/// it on others with nothing to notice it by.
Directive { message: String },
/// An include or ASCII-art target cannot be used. `path` is where the /// An include or ASCII-art target cannot be used. `path` is where the
/// directive actually pointed, resolved, which is the thing an author needs /// directive actually pointed, resolved, which is the thing an author needs
/// to see when a relative target is wrong. /// to see when a relative target is wrong.
@ -76,6 +82,7 @@ impl fmt::Display for Error {
Error::TooDeep { limit } => { Error::TooDeep { limit } => {
write!(f, "document nests more than {limit} levels deep") write!(f, "document nests more than {limit} levels deep")
} }
Error::Directive { message } => write!(f, "{message}"),
Error::Include { path, reason } => { Error::Include { path, reason } => {
write!(f, "include target {} {reason}", path.display()) write!(f, "include target {} {reason}", path.display())
} }
@ -103,4 +110,8 @@ impl Error {
pub(crate) fn config(message: impl Into<String>) -> Self { pub(crate) fn config(message: impl Into<String>) -> Self {
Error::Config { message: message.into() } Error::Config { message: message.into() }
} }
pub(crate) fn directive(message: impl Into<String>) -> Self {
Error::Directive { message: message.into() }
}
} }

View file

@ -1,15 +1,16 @@
//! The one parsed representation every output format consumes. //! The one parsed representation every output format consumes.
//! //!
//! smolweb parses each document twice, with two hand-rolled regex parsers that //! One parse, shared by every format. Parsing per format is how two outputs come
//! share seven identical patterns but disagree on the edges: that is where the //! to disagree about the same document: a directive one understands and another
//! `{.card}` directive leaks into gemtext as literal text, and why the two //! emits as literal text, or two sets of defaults that drift apart. Parsing once
//! libraries carry two divergent sets of defaults. Parsing once into this //! into this structure removes that class of bug rather than its instances.
//! structure removes the class of bug rather than the instances.
//! //!
//! It is a flat block sequence rather than a tree of nodes because that is what //! It is a flat block sequence rather than a tree of nodes because that is what
//! the consumers want: WML packs a linear run of blocks into byte-budgeted //! the consumers want: WML packs a linear run of blocks into byte-budgeted
//! cards, and the text renderers are a fold over blocks. //! cards, and the text renderers are a fold over blocks.
use std::borrow::Cow;
/// A parsed document. /// A parsed document.
#[derive(Debug, Clone, PartialEq, Eq)] #[derive(Debug, Clone, PartialEq, Eq)]
pub struct Doc { pub struct Doc {
@ -64,6 +65,16 @@ pub enum Block {
margin: Option<u16>, margin: Option<u16>,
block: Box<Block>, block: Box<Block>,
}, },
/// A block an `<!-- only ... -->` or `<!-- except ... -->` run restricted to
/// some output formats. Resolved by [`Doc::for_format`] before rendering, so
/// no renderer meets one; it wraps rather than being a field for the same
/// reason `Aligned` does, and the two nest in either order.
Gated {
formats: Vec<String>,
/// `true` for `except`: keep the block for every format *but* these.
negated: bool,
block: Box<Block>,
},
/// Raw block HTML. Kept rather than dropped so XHTML-MP can pass it through /// Raw block HTML. Kept rather than dropped so XHTML-MP can pass it through
/// and the text formats can strip it, instead of the parser deciding. /// and the text formats can strip it, instead of the parser deciding.
Html(String), Html(String),
@ -108,6 +119,19 @@ impl Doc {
out out
} }
/// This document as one output format sees it: a block gated to other
/// formats is dropped and a gate that passes is unwrapped.
///
/// Borrowed unchanged when there are no gates, which is the ordinary page.
/// `first_h1` is kept whatever happens: the title is resolved once per page,
/// so a heading inside a gate still names the page on every format.
pub fn for_format(&self, format: &str) -> Cow<'_, Doc> {
if !any_gated(&self.blocks) {
return Cow::Borrowed(self);
}
Cow::Owned(Doc { blocks: retain(&self.blocks, format), first_h1: self.first_h1.clone() })
}
fn write_plain(inline: &[Inline], out: &mut String) { fn write_plain(inline: &[Inline], out: &mut String) {
for item in inline { for item in inline {
match item { match item {
@ -125,10 +149,149 @@ impl Doc {
} }
} }
/// Whether a gate appears anywhere in a run of blocks, including inside the
/// wrappers a gate can be written in.
fn any_gated(blocks: &[Block]) -> bool {
blocks.iter().any(|block| match block {
Block::Gated { .. } => true,
Block::Aligned { block, .. } => any_gated(std::slice::from_ref(block)),
Block::BlockQuote(inner) => any_gated(inner),
Block::List { items, .. } => items.iter().any(|item| any_gated(item)),
_ => false,
})
}
fn retain(blocks: &[Block], format: &str) -> Vec<Block> {
blocks.iter().filter_map(|block| keep(block, format)).collect()
}
/// One block as `format` sees it, or `None` if a gate excludes it.
fn keep(block: &Block, format: &str) -> Option<Block> {
match block {
Block::Gated { formats, negated, block } => {
let named = formats.iter().any(|id| id == format);
// `only` keeps the formats it names, `except` keeps all the others.
(named != *negated).then(|| keep(block, format)).flatten()
}
Block::Aligned { align, margin, block } => Some(Block::Aligned {
align: *align,
margin: *margin,
block: Box::new(keep(block, format)?),
}),
Block::BlockQuote(inner) => Some(Block::BlockQuote(retain(inner, format))),
Block::List { ordered, start, items } => Some(Block::List {
ordered: *ordered,
start: *start,
items: items.iter().map(|item| retain(item, format)).collect(),
}),
other => Some(other.clone()),
}
}
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
fn para(text: &str) -> Block {
Block::Paragraph(vec![Inline::Text(text.into())])
}
fn gate(formats: &[&str], negated: bool, block: Block) -> Block {
Block::Gated {
formats: formats.iter().map(|id| id.to_string()).collect(),
negated,
block: Box::new(block),
}
}
fn doc(blocks: Vec<Block>) -> Doc {
Doc { blocks, first_h1: None }
}
#[test]
fn a_document_without_gates_is_borrowed_unchanged() {
let source = doc(vec![para("a")]);
assert!(matches!(source.for_format("html"), Cow::Borrowed(_)));
}
#[test]
fn an_only_gate_keeps_the_formats_it_names_and_drops_the_rest() {
let source = doc(vec![gate(&["html", "wml"], false, para("a")), para("b")]);
assert_eq!(source.for_format("html").blocks, vec![para("a"), para("b")]);
assert_eq!(source.for_format("gemtext").blocks, vec![para("b")]);
}
#[test]
fn an_except_gate_drops_the_formats_it_names() {
let source = doc(vec![gate(&["wml"], true, para("a"))]);
assert_eq!(source.for_format("wml").blocks, Vec::new());
assert_eq!(source.for_format("html").blocks, vec![para("a")]);
}
#[test]
fn nested_gates_must_both_pass() {
let source = doc(vec![gate(&["html", "wml"], false, gate(&["wml"], true, para("a")))]);
assert_eq!(source.for_format("html").blocks, vec![para("a")]);
assert_eq!(source.for_format("wml").blocks, Vec::new());
assert_eq!(source.for_format("text").blocks, Vec::new());
}
#[test]
fn a_format_no_gate_names_is_simply_not_named() {
// A build without the wml feature therefore hides an `only wml` run,
// which is the answer that run asked for.
let source = doc(vec![gate(&["wml"], false, para("a"))]);
assert_eq!(source.for_format("gemtext").blocks, Vec::new());
}
#[test]
fn gates_are_resolved_inside_quotes_lists_and_alignment() {
let aligned = Block::Aligned {
align: Align::Center,
margin: None,
block: Box::new(gate(&["html"], false, para("c"))),
};
let source = doc(vec![
Block::BlockQuote(vec![gate(&["html"], false, para("a")), para("b")]),
Block::List {
ordered: false,
start: 1,
items: vec![vec![gate(&["html"], false, para("x"))], vec![para("y")]],
},
aligned,
]);
let kept = source.for_format("html");
assert_eq!(kept.blocks[0], Block::BlockQuote(vec![para("a"), para("b")]));
let Block::List { items, .. } = &kept.blocks[1] else { panic!("expected a list") };
assert_eq!(items, &vec![vec![para("x")], vec![para("y")]]);
assert_eq!(
kept.blocks[2],
Block::Aligned { align: Align::Center, margin: None, block: Box::new(para("c")) }
);
let dropped = source.for_format("text");
assert_eq!(dropped.blocks[0], Block::BlockQuote(vec![para("b")]));
let Block::List { items, .. } = &dropped.blocks[1] else { panic!("expected a list") };
assert_eq!(items, &vec![Vec::new(), vec![para("y")]]);
// An alignment wrapping nothing that survives goes with it.
assert_eq!(dropped.blocks.len(), 2);
}
#[test]
fn the_title_survives_a_gate_around_the_heading() {
// One page, one title, whichever format asks for it.
let source = Doc {
blocks: vec![gate(
&["html"],
false,
Block::Heading { level: 1, inline: vec![Inline::Text("T".into())] },
)],
first_h1: Some("T".into()),
};
assert_eq!(source.for_format("gemtext").first_h1.as_deref(), Some("T"));
}
#[test] #[test]
fn plain_text_flattens_markup_and_keeps_link_labels() { fn plain_text_flattens_markup_and_keeps_link_labels() {
let inline = vec![ let inline = vec![

View file

@ -1,9 +1,9 @@
//! Media types for files served byte for byte. //! Media types for files served byte for byte.
//! //!
//! A fixed table rather than a system lookup: Python's `mimetypes` consults //! A fixed table rather than a system lookup. A lookup reads `/etc/mime.types`
//! `/etc/mime.types` where it exists, so smolweb's `Content-Type` for the same //! where it exists, so the same file gets a different `Content-Type` depending on
//! file differs between hosts. Determinism is worth more here than coverage of //! the host. Determinism is worth more here than coverage of the long tail, and
//! the long tail, and an unknown extension has a correct answer anyway. //! an unknown extension has a correct answer anyway.
use std::path::Path; use std::path::Path;

View file

@ -427,8 +427,6 @@ mod tests {
#[test] #[test]
fn parses_setext_headings() { fn parses_setext_headings() {
// wapdown's parser handles these and md2txt's does not, so unifying the
// two parsers gains them for the text formats.
assert_eq!( assert_eq!(
blocks("Title\n=====\n"), blocks("Title\n=====\n"),
vec![Block::Heading { level: 1, inline: text("Title") }] vec![Block::Heading { level: 1, inline: text("Title") }]

View file

@ -82,8 +82,8 @@ fn encode_path(path: &str) -> String {
/// Decode `%XX` escapes, leaving an invalid escape as the literal text it is. /// Decode `%XX` escapes, leaving an invalid escape as the literal text it is.
/// ///
/// Bytes that do not form valid UTF-8 become U+FFFD, which matches no filename, /// Bytes that do not form valid UTF-8 become U+FFFD, which matches no filename,
/// so a malformed target resolves to nothing rather than erroring. That matches /// so a malformed target resolves to nothing rather than erroring. A test pins
/// Python's lossy `unquote` and is pinned by a test. /// that leniency, so it is not later tightened into an error.
fn percent_decode(raw: &str) -> String { fn percent_decode(raw: &str) -> String {
let bytes = raw.as_bytes(); let bytes = raw.as_bytes();
let mut out = Vec::with_capacity(bytes.len()); let mut out = Vec::with_capacity(bytes.len());
@ -145,9 +145,8 @@ mod tests {
#[test] #[test]
fn clamps_traversal_at_the_root() { fn clamps_traversal_at_the_root() {
// Ported from smolweb's TestPathTraversal: these must never reach above // These must never reach above the root, and since nothing is mounted at
// the root, and since nothing is mounted at the clamped path they // the clamped path they resolve to a path that simply does not exist.
// resolve to a path that simply does not exist.
assert_eq!(clean("/../../etc/passwd"), "etc/passwd"); assert_eq!(clean("/../../etc/passwd"), "etc/passwd");
assert_eq!(clean("/../../../../../../etc/passwd"), "etc/passwd"); assert_eq!(clean("/../../../../../../etc/passwd"), "etc/passwd");
assert_eq!(clean("/foo/../../etc/passwd"), "etc/passwd"); assert_eq!(clean("/foo/../../etc/passwd"), "etc/passwd");
@ -159,7 +158,7 @@ mod tests {
// before the clamp runs, or it would be treated as a literal segment. // before the clamp runs, or it would be treated as a literal segment.
assert_eq!(clean("/%2e%2e/etc/passwd"), "etc/passwd"); assert_eq!(clean("/%2e%2e/etc/passwd"), "etc/passwd");
assert_eq!(clean("/%2E%2E/etc/passwd"), "etc/passwd"); assert_eq!(clean("/%2E%2E/etc/passwd"), "etc/passwd");
// An encoded separator becomes a separator, as Python's unquote does. // An encoded separator becomes a real one, so normalising then sees it.
assert_eq!(clean("/dir%2fpage"), "dir/page"); assert_eq!(clean("/dir%2fpage"), "dir/page");
assert_eq!(clean("/hello%20world"), "hello world"); assert_eq!(clean("/hello%20world"), "hello world");
} }

View file

@ -7,10 +7,10 @@
//! [`crate::directives`]. //! [`crate::directives`].
//! //!
//! Together with that module this is the only part of the pipeline that opens //! Together with that module this is the only part of the pipeline that opens
//! files, which gives the root-containment check exactly one home. That closes //! files, which gives the root-containment check exactly one home. An include
//! the traversal smolweb has: md2txt resolves an include target and checks only //! target is canonicalised and required to be inside the root, so a target of
//! that it exists, so `{.include ../../../../etc/passwd}` in any served document //! `../../../../etc/passwd` resolves to nothing instead of being read: checking
//! reads and emits that file. //! only that a target exists is what makes includes a traversal.
use std::collections::BTreeSet; use std::collections::BTreeSet;
use std::fs; use std::fs;
@ -24,7 +24,7 @@ use crate::error::{Error, IncludeReason};
const MAX_DEPTH: usize = 16; const MAX_DEPTH: usize = 16;
/// Caps on the expanded result. The cycle set is per-*stack*, so a diamond — /// Caps on the expanded result. The cycle set is per-*stack*, so a diamond —
/// `a` includes `b` and `c`, both include `d` — fans out exponentially without /// `a` includes `b` and `c`, both include `d` — fans out exponentially without
/// ever repeating a file on one path. smolweb has nothing that stops this. /// ever repeating a file on one path, so only these caps stop it.
const MAX_LINES: usize = 200_000; const MAX_LINES: usize = 200_000;
const MAX_BYTES: usize = 8 * 1024 * 1024; const MAX_BYTES: usize = 8 * 1024 * 1024;
@ -159,7 +159,7 @@ mod tests {
assert_eq!(include_target("see ![[notes.md]] there"), None); assert_eq!(include_target("see ![[notes.md]] there"), None);
assert_eq!(include_target("![[]]"), None); assert_eq!(include_target("![[]]"), None);
assert_eq!(include_target("![[unterminated"), None); assert_eq!(include_target("![[unterminated"), None);
// The directive smolweb also accepted is gone: one spelling, not two. // The brace-directive form is not accepted: one spelling, not two.
assert_eq!(include_target("{.include notes.md}"), None); assert_eq!(include_target("{.include notes.md}"), None);
} }
@ -263,8 +263,8 @@ mod tests {
#[test] #[test]
fn a_diamond_fan_out_is_stopped_by_the_size_cap() { fn a_diamond_fan_out_is_stopped_by_the_size_cap() {
// Each level doubles and no file repeats on any single path, so neither // Each level doubles and no file repeats on any single path, so neither
// the cycle set nor the depth cap catches it. smolweb expands this until // the cycle set nor the depth cap catches it, which leaves the byte cap
// it runs out of memory. // as the only thing that stops it.
let tree = Tree::new(); let tree = Tree::new();
tree.write("leaf.md", &"filler line\n".repeat(64)); tree.write("leaf.md", &"filler line\n".repeat(64));
let mut previous = "leaf.md".to_string(); let mut previous = "leaf.md".to_string();

View file

@ -3,12 +3,12 @@
//! A format is a crate implementing [`Renderer`], registered at startup behind a //! A format is a crate implementing [`Renderer`], registered at startup behind a
//! cargo feature. The trait takes a parsed [`Doc`] rather than source text so //! cargo feature. The trait takes a parsed [`Doc`] rather than source text so
//! that every format reads one parse: gemtext's `=>` link catalogue and the text //! that every format reads one parse: gemtext's `=>` link catalogue and the text
//! formats' `[n]` references must agree about link identity and order, and in //! formats' `[n]` references must agree about link identity and order, which
//! smolweb they can disagree because each library re-parses. //! cannot be relied on when each format parses the source for itself.
//! //!
//! A renderer must not open files or sockets. Includes and art are already //! A renderer must not open files or sockets. Includes, art and format gates are
//! resolved by the time it runs, which is what keeps the root-containment check //! already resolved by the time it runs, which is what keeps the
//! in one place rather than in every format. //! root-containment check in one place rather than in every format.
use std::collections::BTreeMap; use std::collections::BTreeMap;
use std::sync::Arc; use std::sync::Arc;
@ -133,7 +133,10 @@ impl Registry {
})?; })?;
let ctx = let ctx =
RenderCtx { url, title: &page.title, settings, width: renderer.default_width() }; RenderCtx { url, title: &page.title, settings, width: renderer.default_width() };
let rendered = renderer.render(doc, &ctx)?; // Format gates are resolved here, so no renderer meets one and a
// gated run costs nothing extra in the cache: bodies are already
// kept per format.
let rendered = renderer.render(&doc.for_format(id), &ctx)?;
for part in rendered.parts { for part in rendered.parts {
page.parts.insert((id.clone(), part.slug), part.body); page.parts.insert((id.clone(), part.slug), part.body);
} }
@ -202,9 +205,12 @@ mod tests {
self.width self.width
} }
fn render(&self, _doc: &Doc, ctx: &RenderCtx<'_>) -> Result<Rendered, Error> { fn render(&self, doc: &Doc, ctx: &RenderCtx<'_>) -> Result<Rendered, Error> {
Ok(Rendered::body( Ok(Rendered::body(
format!("{} {} {:?} {}", self.id, ctx.url, ctx.width, ctx.title).into_bytes(), // The blocks are printed too, so what a format was handed —
// after gates — is visible in the body.
format!("{} {} {:?} {} {:?}", self.id, ctx.url, ctx.width, ctx.title, doc.blocks)
.into_bytes(),
)) ))
} }
} }
@ -264,6 +270,23 @@ mod tests {
assert_eq!(from_stem.unwrap().title, "stem"); assert_eq!(from_stem.unwrap().title, "stem");
} }
#[test]
fn a_gated_block_reaches_only_the_formats_it_names() {
let gated = Doc {
blocks: vec![Block::Gated {
formats: vec!["one".to_string()],
negated: false,
block: Box::new(Block::Paragraph(vec![Inline::Text("secret".into())])),
}],
first_h1: None,
};
let formats = vec!["one".to_string(), "two".to_string()];
let page = registry().page(&formats, &gated, "/x", &PageSettings::default(), "x").unwrap();
// The stub prints the blocks it was handed, so presence is visible.
assert!(String::from_utf8_lossy(page.body("one").unwrap()).contains("secret"));
assert!(!String::from_utf8_lossy(page.body("two").unwrap()).contains("secret"));
}
#[test] #[test]
fn an_unknown_format_is_a_bug_not_bad_input() { fn an_unknown_format_is_a_bug_not_bad_input() {
let err = registry() let err = registry()

View file

@ -267,8 +267,8 @@ mod tests {
Site::new(root, registry(), vec!["stub".to_string()]).unwrap() Site::new(root, registry(), vec!["stub".to_string()]).unwrap()
} }
/// Mirrors smolweb's `tests/conftest.py` fixture, so its assertions port /// One of each kind of thing a request can land on, shared by the resolution
/// across directly. /// tests below.
fn fixture() -> (tempfile::TempDir, Site) { fn fixture() -> (tempfile::TempDir, Site) {
let dir = tempfile::tempdir().unwrap(); let dir = tempfile::tempdir().unwrap();
let root = dir.path(); let root = dir.path();
@ -449,9 +449,9 @@ mod tests {
#[test] #[test]
fn a_bare_unresolvable_segment_resolves_to_nothing() { fn a_bare_unresolvable_segment_resolves_to_nothing() {
// Regression carried over from smolweb: the Python reached this path // The obvious implementation splits on the last `/`, where a bare
// through `rpartition("/")`, where a bare top-level segment yields an // top-level segment yields an empty parent that must not then be read as
// empty parent that must not be read as the root index. // the root index.
let (_dir, site) = fixture(); let (_dir, site) = fixture();
assert_not_found(&site, "/totally-unresolvable-segment"); assert_not_found(&site, "/totally-unresolvable-segment");
} }
@ -568,9 +568,9 @@ mod tests {
#[cfg(test)] #[cfg(test)]
mod part_tests { mod part_tests {
//! Ported from smolweb's `TestWmlCardUrls`. A stub renderer stands in for a //! A stub renderer stands in for a paginating format, so these rules are
//! paginating format, so these rules are tested without the WML crate: core //! tested without the WML crate: core does not know which formats paginate,
//! does not know which formats paginate, which is the point. //! which is the point.
use std::fs; use std::fs;

37
docker/itsybitsy.toml Normal file
View file

@ -0,0 +1,37 @@
version = 1
[site.demo]
root = "./content"
hosts = ["localhost", "127.0.0.1"]
[listener.web]
protocol = "http"
bind = "0.0.0.0:8080"
formats = ["wml", "xhtmlmp", "html"]
default_site = "demo"
[listener.spartan]
protocol = "spartan"
bind = "0.0.0.0:3000"
formats = ["gemtext"]
default_site = "demo"
[listener.nex]
protocol = "nex"
bind = "0.0.0.0:1900"
formats = ["text"]
site = "demo"
# Plaintext: a TLS terminator in front of the container owns 1965 and forwards
# here, as it does when deployed bare.
[listener.gemini]
protocol = "gemini"
bind = "0.0.0.0:1965"
formats = ["gemtext"]
default_site = "demo"
[listener.gopher]
protocol = "gopher"
bind = "0.0.0.0:7070"
formats = ["text"]
site = "demo"

View file

@ -143,6 +143,9 @@ impl Writer {
// Alignment and pagination have no expression here. Unwrapping keeps // Alignment and pagination have no expression here. Unwrapping keeps
// the content; the presentation is simply not available. // the content; the presentation is simply not available.
Block::Aligned { block, .. } => self.block(block), Block::Aligned { block, .. } => self.block(block),
// Gates are filtered out before rendering; keeping the content is
// the harmless reading if one ever arrives here.
Block::Gated { block, .. } => self.block(block),
Block::CardBreak { .. } => {} Block::CardBreak { .. } => {}
// Raw markup would be shown literally, which is worse than omitting it. // Raw markup would be shown literally, which is worse than omitting it.
Block::Html(_) => {} Block::Html(_) => {}
@ -205,7 +208,13 @@ fn is_only_links(inline: &[Inline]) -> bool {
match item { match item {
Inline::Link { .. } | Inline::Image { .. } => saw_link = true, Inline::Link { .. } | Inline::Image { .. } => saw_link = true,
Inline::Text(text) if text.trim().is_empty() => {} Inline::Text(text) if text.trim().is_empty() => {}
Inline::SoftBreak => {} // Either break kind. A hard break between links is still a run of
// nothing but links, and a generator has good reason to use one:
// the formats that lay links out inline need <br/> to stop a menu
// running together on one line, and refusing it here would make
// the paragraph print its flattened text as well as the => lines,
// listing every entry twice.
Inline::SoftBreak | Inline::HardBreak => {}
_ => return false, _ => return false,
} }
} }
@ -300,6 +309,16 @@ mod tests {
use super::tests_support::render; use super::tests_support::render;
use super::*; use super::*;
#[test]
fn a_hard_break_between_links_still_lifts_them() {
// Generators emit hard breaks so the inline formats put each link on
// its own line. gemtext must not read that as prose and print the
// flattened text alongside the => lines, which would list every entry
// twice.
let out = render("[one](/a) \n[two](/b)\n");
assert_eq!(out, "=> /a one\n=> /b two\n");
}
#[test] #[test]
fn headings_flatten_onto_three_levels() { fn headings_flatten_onto_three_levels() {
assert_eq!(render("# a\n## b\n### c\n#### d\n"), "# a\n\n## b\n\n### c\n\n### d\n"); assert_eq!(render("# a\n## b\n### c\n#### d\n"), "# a\n\n## b\n\n### c\n\n### d\n");

View file

@ -2,9 +2,7 @@
//! //!
//! Plain text has no markup to carry emphasis, so it is dropped and the words //! Plain text has no markup to carry emphasis, so it is dropped and the words
//! kept. A link becomes `label (url)`: self-contained, and readable without //! kept. A link becomes `label (url)`: self-contained, and readable without
//! scrolling to a reference list somewhere else. md2txt's `text` renderer emits //! scrolling to a reference list somewhere else.
//! numbered markers instead but never writes the list they point at, so the
//! numbers lead nowhere.
use itsybitsy_core::ir::{Doc, Inline}; use itsybitsy_core::ir::{Doc, Inline};

View file

@ -122,6 +122,9 @@ impl<'a> Layout<'a> {
prefix, prefix,
Placement { align: Some(*align), inset: margin.unwrap_or(0) as usize }, Placement { align: Some(*align), inset: margin.unwrap_or(0) as usize },
), ),
// Gates are filtered out before rendering; keeping the content is
// the harmless reading if one ever arrives here.
Block::Gated { block, .. } => self.block(block, prefix, place),
// Nothing paginates here, so a divider has nothing to divide. // Nothing paginates here, so a divider has nothing to divide.
Block::CardBreak { .. } => {} Block::CardBreak { .. } => {}
Block::Html(_) => {} Block::Html(_) => {}

View file

@ -1,10 +1,10 @@
//! Fixed-width plain text, for Nex and later Gopher. //! Fixed-width plain text, for Nex and later Gopher.
//! //!
//! One renderer, not two. md2txt ships a `text` and a `nex` renderer that differ //! One renderer, not two: whether a heading gets a FIGlet banner and how a link
//! in exactly two things — whether headings get FIGlet banners, and whether links //! is written are both configuration, so Nex and Gopher are this renderer with
//! are inlined or numbered — and both of those are now configuration. Its //! different settings rather than renderers of their own. A link is inlined as
//! numbered form never writes the reference list its numbers point at, so the //! `label (url)` rather than numbered, because a numbered marker is only useful
//! inline form is the only one that works and is the default here. //! with a reference list to point at.
//! //!
//! Unlike gemtext this wraps, because Nex and Gopher clients do not. //! Unlike gemtext this wraps, because Nex and Gopher clients do not.
@ -115,7 +115,7 @@ mod tests {
#[test] #[test]
fn one_blank_line_separates_blocks_by_default() { fn one_blank_line_separates_blocks_by_default() {
// md2txt emits two, which reads as double-spaced throughout. // Two would read as double-spaced throughout.
assert_eq!(render("a\n\nb\n"), "a\n\nb\n"); assert_eq!(render("a\n\nb\n"), "a\n\nb\n");
} }
@ -184,7 +184,6 @@ mod tests {
#[test] #[test]
fn tables_are_rendered_as_aligned_columns() { fn tables_are_rendered_as_aligned_columns() {
// md2txt drops tables entirely; this is the flaw that fixes.
assert_eq!( assert_eq!(
render("| Format | Port |\n| --- | --- |\n| Nex | 1900 |\n"), render("| Format | Port |\n| --- | --- |\n| Nex | 1900 |\n"),
"Format Port\n------ ----\nNex 1900\n" "Format Port\n------ ----\nNex 1900\n"

View file

@ -77,6 +77,9 @@ fn block_markup(block: &Block, settings: &PageSettings, out: &mut String) {
Block::Rule => out.push_str("<p>---</p>\n"), Block::Rule => out.push_str("<p>---</p>\n"),
// Alignment is a fixed-width concern; a handset lays out its own screen. // Alignment is a fixed-width concern; a handset lays out its own screen.
Block::Aligned { block, .. } => block_markup(block, settings, out), Block::Aligned { block, .. } => block_markup(block, settings, out),
// Gates are filtered out before rendering; keeping the content is the
// harmless reading if one ever arrives here.
Block::Gated { block, .. } => block_markup(block, settings, out),
// Pagination here is the deck's own job, driven by the byte budget. // Pagination here is the deck's own job, driven by the byte budget.
Block::CardBreak { .. } => {} Block::CardBreak { .. } => {}
// Raw HTML is not WML and would not parse on a handset. // Raw HTML is not WML and would not parse on a handset.
@ -184,7 +187,6 @@ mod tests {
#[test] #[test]
fn tables_are_rendered_because_wml_has_them() { fn tables_are_rendered_because_wml_has_them() {
// wapdown's own parser produces no tables, so this is new output.
assert_eq!( assert_eq!(
render("| a | b |\n| --- | --- |\n| 1 | 2 |\n"), render("| a | b |\n| --- | --- |\n| 1 | 2 |\n"),
"<table columns=\"2\">\n<tr><td>a</td><td>b</td></tr>\n<tr><td>1</td><td>2</td></tr>\n</table>\n" "<table columns=\"2\">\n<tr><td>a</td><td>b</td></tr>\n<tr><td>1</td><td>2</td></tr>\n</table>\n"

View file

@ -112,6 +112,9 @@ fn self_block(block: &Block, out: &mut String) {
// Alignment reaches only the fixed-width text formats. Pagination is a // Alignment reaches only the fixed-width text formats. Pagination is a
// WML concern; a browser scrolls one document. // WML concern; a browser scrolls one document.
Block::Aligned { block, .. } => self_block(block, out), Block::Aligned { block, .. } => self_block(block, out),
// Gates are filtered out before rendering; keeping the content is the
// harmless reading if one ever arrives here.
Block::Gated { block, .. } => self_block(block, out),
Block::CardBreak { .. } => {} Block::CardBreak { .. } => {}
} }
} }