This commit is contained in:
randogoth 2025-10-20 21:58:05 +03:00
parent 2db4eab94c
commit ea43375cb4
5 changed files with 415 additions and 8 deletions

View file

@ -19,6 +19,7 @@ uv run mambler.py --title "Your Book Title" --codepage 437 path/to/index.md outp
- `--title` is optional; when provided the value is embedded in the AMB archive header (truncated to 64 ASCII bytes). - `--title` is optional; when provided the value is embedded in the AMB archive header (truncated to 64 ASCII bytes).
- `--codepage` controls the 8-bit encoding used for every AMA article (default: `437`). Any character that cannot be expressed in the chosen codepage aborts the build with a helpful error so you can pick a better fit. - `--codepage` controls the 8-bit encoding used for every AMA article (default: `437`). Any character that cannot be expressed in the chosen codepage aborts the build with a helpful error so you can pick a better fit.
- If any emitted byte lives in the 0x800xFF range, `mambler` automatically writes a companion `UNICODE.MAP` file describing the high-half character mapping, mirroring the recommendation in the AMA/AMB specification. - If any emitted byte lives in the 0x800xFF range, `mambler` automatically writes a companion `UNICODE.MAP` file describing the high-half character mapping, mirroring the recommendation in the AMA/AMB specification.
- Words of length 217 are indexed into `DICT.IDX` so readers can offer fast full-text search. The index is omitted if it would overflow the 64KiB LoW data limit mandated by the spec.
- The command prints the path of the generated AMB file on success. - The command prints the path of the generated AMB file on success.
### Development Notes ### Development Notes

259
format.txt Normal file
View file

@ -0,0 +1,259 @@
==== AMB FORMAT SPECIFICATION ====
last updated: 2025-08-26
The latest version of this file can be found on the AMB project's homepage:
<http://mateusz.fr/amb/>
An AMB file (Ancient Machine Book) is an extremely lightweight file format
meant to store any kind of hypertext documentation that may be comfortably
viewed even on the most ancient PCs: technical manuals, books, etc. Think of
it as a retro equivalent of a *.CHM help file. The AMB format is designed to
allow for some limited formatting, support internal links and require very
little processing power to read, so a reader may be run even on the oldest
IBM PC. The format also strives for simplicity of implementation.
Table Of Contents:
* The AMB container
* Title
* AMA format
* Codepage encoding
* Index data
* Rationale
==============================================================================
THE AMB CONTAINER
The AMB file is a container - one could say it is a very simplistic archive
format. It starts with a 4-bytes format signature (magic value) "AMB1". Then
comes a 2-bytes number that tells how many files are present in the container,
followed by the list of all files: each file is described by a file entry.
All values are little-endian.
offset
0 format signature: "AMB1"
4 files count (16-bit value)
6 FILE ENTRY #1
FILE ENTRY #2
FILE ENTRY #3
....
DATA
Each file entry is a 20-bytes structure:
offset
0 filename, 12 characters, zero-padded ("FILE.EXT\0\0\0\0")
12 offset where this file starts (32 bits)
16 file length, in bytes (16 bits)
18 BSD sum (16-bit) of the file
The AMB archive is expected to contain a set of AMA (Ancient Machine Article)
files, and optionally a title file, an index dictionary and a codepage map.
AMA files may be compressed with the MVCOMP algorithm, in which case they are
named with the "*.AMC" extension.
An AMB archive must contain at least one article file named either "index.ama"
or "index.amc" - this is the first file that an AMB reader will try loading.
Note: Names of files contained in an AMB archive are to be processed in a case
insensitive way and must be composed exclusively of 7-bit characters.
==============================================================================
DOCUMENT TITLE
The AMB title is a string that may be displayed as the document's main title.
To set such title, the AMB archive has to contain a file named simply 'title'
that would contain the text. The title string should not be longer than
64 characters, anything longer might be truncated by the reader.
The title of the document is expected to be encoded with the same codepage as
all the articles. See codepage encoding.
==============================================================================
AMA FORMAT
The AMA format is a text-based file format. For guaranteed interoperability
with old machines, its maximum allowed size is 65535 bytes (ie. 2^16 - 1).
Larger contents must be segmented into a set of two or more AMA articles.
An AMB reader must display content with a 78-characters width, hence an AMA
article must not contain any line longer than 78 displayable characters. Lines
longer than this limit may be truncated by the client reader.
AMA articles may contain control codes. A control code is a characters pair,
where the first is a percent (%) character. Possible control codes:
%t normal text follows (default state)
%h heading follows
%l link follows (filename ended by a ':', followed by a description)
%! notice/warning follows
%b boring text follows (usually displayed grey on grey)
%% display a percent character (%)
It is important to note that the current text mode is reset to %t at the end
of every line, hence there is no need to prefix a line of text with %t.
Line endings may be either LF or CR/LF. The former is recommended, as it is
more compact.
TAB control codes (ASCII decimal value 9) are NOT allowed in AMA files.
Whenever an external URL appears in an AMA file (for example a link to a web
page, to a ftp resource or to a gopher hole) it is encouraged to be enclosed
between <> characters. Example: <http://mateusz.fr/amb>. This is a typesetting
recommendation based on RFC 3986, it is not part of the AMA specification.
Following it would, however, make it much easier for modern AMB readers to
detect such links automatically and make them clickable.
==============================================================================
CODEPAGE ENCODING
Since ancient computers are displaying text as 8-bit characters due to the
design of early video adapters, AMA files are expected to contain 8-bit text
as well. The exact codepage is unspecified by this format definition and
depends on the document's target audience.
To ease displaying of AMB books on modern (unicode-enabled) platforms, any AMB
file that contains non-7-bit characters SHOULD also contain a file named
"unicode.map". This file contains a sequence of 128 16-bit values, mapping
bytes of the range 128..255 into unicode datapoint values. Such file can be
readily output by the utf8tocp program <http://mateusz.fr/utf8tocp/>.
==============================================================================
INDEX DATA
On top of AMA files, the AMB archive may contain a file named DICT.IDX. This
file, if it exists, provides indexing metadata to allow the client to perform
fast and efficient full-text searches across the AMB book.
The index file contains a hash table: a serie of 256 16-bit indexes, where
each index points to a region of the index structure that contains a list of
words (LoW). The index (0..255) itself is an 8 bits hash based on the length
of the word and its characters. The checksum is made of two nibbles: LC.
The high nibble (L) is the length of the word minus 2, while the low nibble
(C) is a simple checksum of all the word's characters XORed together. This
algorithm can be formalized as follows:
((wordlen - 2) << 4) | ((a & 15) XOR (b & 15) XOR (...))
For example, the word "Disk" would end up being indexed under value 0x25,
because:
((4 - 2) << 4) | ((D & 15) XOR (i & 15) XOR (s & 15) XOR (k & 15))
translates to: (2 << 4) | (4 XOR 9 XOR 3 XOR 11)
which leads to: 32 | 5
resulting in: 37 = 0x25
After the index we can find the pointer to the words list. A pointer is a 16
bits file offset from the index structure start.
It needs to be noted that words of less than 2 characters and more than 17
characters cannot be indexed. The presented algorithm has also the interesting
side-effect of indexing low and high caps of the ranges a..z and A..Z
identically. An important limitation is the fact that the list of words (LoW)
is restricted by the 16-bit addressing offset, which means that all LoWs must
start at an offset within the first 64 KiB of the file.
Now that we know the offset at which our LoW starts, we can read the words.
First go to the offset, and read a single 16 bits word. Its value contains the
number of words in the list. Then, read the words one after another (note that
all words in the list have the same length, and you know this length already).
Words are always written in lower case characters. Each word is followed by a
1-byte value that tells how many files the word has been found in. Then, that
many 32-bit file identifiers follow.
index format:
* List of words
xx number of words in the list
? word
x how many files the word is present in
xxxx file identifier 1
xxxx file identifier 2
...
xxxx file identifier n
(other 255 lists of words follow)
* hash table
xx offset of the LoW for words that match hash 0x00
xx offset of the LoW for words that match hash 0x01
...
xx offset of the LoW for words that match hash 0xff
==============================================================================
RATIONALE
The AMB format is, by design, burdened by several limitations. These
limitations might be misunderstood as shortcomings, while in essence the AMB
format's primary objective is to stay as primitive as possible - so it is easy
(and fast) for software to parse and display. Below are listed some of these
limitations, with explanations about the reasons that led to them.
* Line length limited to 78 characters
The hard-coded limit of 78 chars is meant to ensure that the reader will
not have to worry about line wrapping, which highly simplifies the reader's
code thus allowing for faster processing and minimizing potential bugs. It
is also meant to allow the content creator to design his screens in a
deterministic way - that is, without any risk that his semigraphic tables,
ASCII drawings or overall screen disposition will be broken by a reader
that attempts to rewrap the text at an unpredictable width.
The 80-columns width was ubiquitous since the early 80' and seems to be a
reasonable baseline expectation, and a 78-characters limit allows the
reader to use two columns for its own needs (vertical cursor, border, etc).
* No control over style (colors) applied to text
The AMB format defines a set of semantic tags (like "%h" for "heading"). It
does not allow control over the exact colors or attributes that will be
used by the output device to render the document. This is designed on
purpose: AMB documents should be displayable also on monochrome devices.
There may also be devices that allow for text attributes like "underlined",
"bold", etc - it is up to the AMB reader to make sure the semantic tags are
translated into colors/shades/attributes combinations that are nicely
rendered on the target hardware.
* Article size limit of 64 KiB
A single article (AMA file) is limited to a maximum length of 64 KiB (minus
one byte). This limitation makes it easier for MS-DOS readers to load the
content: in real-mode Intel memory models, a single memory segment is
addressable via 16-bit offsets, hence processing content larger than 64 KiB
becomes tricky, as it involves crossing memory segment boundaries, or
relying on some kludges like "huge" memory pointers (slow), or dynamically
reloading parts of the file from disk (very slow). 64 KiB still allows for
more than 30 pages of 80x25 packed text, which should be more than enough
even for very complex subjects (and larger contents should simply be
dispatched into two or more different articles, which can only be
beneficial for readability).
* Maximum number of 65535 articles
An AMB book may contain up to 65535 articles and not a single more, because
the number of articles is written as a 16-bit integer in the file's header.
This allows AMB software to use 16-bit integers when addressing the
articles, which is very convenient (and fast) for platforms with 16-bit
CPUs. And honestly - is that really a limitation? Even the entire Bible has
"only" 1189 chapters, or 31103 verses.
* Short filenames + low-ascii characters only
Filenames inside an AMB container are limited to 12 (8+3) characters so
an AMB container can be unpacked on an old MS-DOS system.
The filenames must contain only low-ascii (7-bit) characters -- for two
reasons: so it is possible to unpack an AMB container on any filesystem,
independently of the codepage said filesystem relies on, and to make it
possible to reliably perform case-insensitive matching of filenames.
==============================================================================

BIN
long.amb

Binary file not shown.

View file

@ -5,10 +5,11 @@ import argparse
import codecs import codecs
import re import re
import struct import struct
import sys
from collections import deque from collections import deque
from dataclasses import dataclass from dataclasses import dataclass
from pathlib import Path from pathlib import Path
from typing import Callable, Dict, Iterable, List, Tuple from typing import Callable, Dict, Iterable, List, Optional, Tuple
from md2txt import convert_markdown from md2txt import convert_markdown
from md2txt.conversion.core import parse_frontmatter from md2txt.conversion.core import parse_frontmatter
@ -249,6 +250,114 @@ def _has_high_bit(data: bytes) -> bool:
return any(byte >= 0x80 for byte in data) return any(byte >= 0x80 for byte in data)
def compute_file_offsets(files: List[Tuple[str, bytes]], include_dict: bool) -> Dict[str, int]:
total_files = len(files) + (1 if include_dict else 0)
offset = 6 + 20 * total_files
mapping: Dict[str, int] = {}
for name, data in files:
mapping[name.upper()] = offset
offset += len(data)
return mapping
def build_dict_index(
word_index: Dict[str, set[str]],
file_offsets: Dict[str, int],
codepage: CodepageInfo,
) -> Optional[bytes]:
bucket_words: Dict[int, Dict[str, Tuple[bytes, List[int]]]] = {}
for word, filenames in word_index.items():
try:
encoded_word = codepage.encode(word)
except UnicodeEncodeError:
continue
length = len(encoded_word)
if not (2 <= length <= 17):
continue
bucket = compute_word_hash(encoded_word)
file_ids = sorted({file_offsets[name.upper()] for name in filenames if name.upper() in file_offsets})
if not file_ids:
continue
bucket_map = bucket_words.setdefault(bucket, {})
bucket_map[word] = (encoded_word, file_ids)
if not bucket_words:
return None
lows = bytearray()
offsets: List[int] = []
for bucket in range(256):
offsets.append(len(lows))
entries = bucket_words.get(bucket)
if not entries:
lows.extend(struct.pack("<H", 0))
continue
sorted_entries = sorted(entries.items(), key=lambda item: item[0])
word_length = len(sorted_entries[0][1][0])
lows.extend(struct.pack("<H", len(sorted_entries)))
for word, (encoded_word, file_ids) in sorted_entries:
if len(encoded_word) != word_length:
raise ValueError(f"Inconsistent word length for hash bucket 0x{bucket:02x}.")
if len(file_ids) > 255:
raise ValueError(f"Word '{word}' appears in more than 255 files, cannot encode index.")
lows.extend(encoded_word)
lows.append(len(file_ids))
for file_id in file_ids:
lows.extend(struct.pack("<I", file_id))
if len(lows) >= 0x10000:
raise ValueError("Generated DICT.IDX exceeds 64 KiB limit for word lists.")
hash_table = bytearray()
for offset in offsets:
hash_table.extend(struct.pack("<H", offset))
return bytes(lows + hash_table)
def compute_word_hash(encoded_word: bytes) -> int:
length = len(encoded_word)
checksum = 0
for byte in encoded_word:
checksum ^= (byte & 0x0F)
return ((length - 2) << 4) | (checksum & 0x0F)
WORD_MIN_LENGTH = 2
WORD_MAX_LENGTH = 17
def extract_words(lines: List[str]) -> set[str]:
words: set[str] = set()
for line in lines:
stripped = _strip_control_codes(line)
buffer: List[str] = []
for char in stripped:
if char.isalnum():
buffer.append(char.lower())
continue
if len(buffer) >= WORD_MIN_LENGTH:
word = "".join(buffer)
if WORD_MIN_LENGTH <= len(word) <= WORD_MAX_LENGTH:
words.add(word)
buffer = []
if len(buffer) >= WORD_MIN_LENGTH:
word = "".join(buffer)
if WORD_MIN_LENGTH <= len(word) <= WORD_MAX_LENGTH:
words.add(word)
return words
def _strip_control_codes(line: str) -> str:
result = line
result = re.sub(r"%l[^:]+:", "", result)
result = result.replace("%t", "").replace("%!", "").replace("%b", "").replace("%h", "")
result = result.replace("%%", "%")
return result
@dataclass @dataclass
class Article: class Article:
source: Path source: Path
@ -287,8 +396,31 @@ def main(argv: Iterable[str] | None = None) -> int:
def build_amb(root_markdown: Path, title: str | None, codepage: CodepageInfo) -> bytes: def build_amb(root_markdown: Path, title: str | None, codepage: CodepageInfo) -> bytes:
articles = collect_articles(root_markdown) articles = collect_articles(root_markdown)
ama_contents = render_articles(articles, codepage) ama_contents, word_index = render_articles(articles, codepage)
files = assemble_files(ama_contents, title, codepage) base_files = assemble_files(ama_contents, title, codepage)
file_offsets = compute_file_offsets(base_files, include_dict=False)
try:
dict_bytes = build_dict_index(word_index, file_offsets, codepage)
except ValueError as exc:
print(f"[mambler] Skipping dictionary index: {exc}", file=sys.stderr)
dict_bytes = None
if dict_bytes is None:
files = base_files
else:
adjusted_offsets = compute_file_offsets(base_files, include_dict=True)
try:
dict_bytes_adjusted = build_dict_index(word_index, adjusted_offsets, codepage)
except ValueError as exc:
print(f"[mambler] Skipping dictionary index: {exc}", file=sys.stderr)
files = base_files
else:
if dict_bytes_adjusted is None:
files = base_files
else:
files = base_files + [("DICT.IDX", dict_bytes_adjusted)]
return pack_amb(files) return pack_amb(files)
@ -351,8 +483,9 @@ def assign_ama_name(stem: str, existing: set[str]) -> str:
return name return name
def render_articles(articles: Dict[Path, Article], codepage: CodepageInfo) -> Dict[str, List[str]]: def render_articles(articles: Dict[Path, Article], codepage: CodepageInfo) -> Tuple[Dict[str, List[str]], Dict[str, set[str]]]:
rendered: Dict[str, List[str]] = {} rendered: Dict[str, List[str]] = {}
word_index: Dict[str, set[str]] = {}
for path, article in articles.items(): for path, article in articles.items():
content = path.read_text(encoding="utf-8") content = path.read_text(encoding="utf-8")
@ -366,8 +499,20 @@ def render_articles(articles: Dict[Path, Article], codepage: CodepageInfo) -> Di
renderer_name="ama", renderer_name="ama",
) )
split_articles = split_article(article.ama_name, ama_lines, codepage) split_articles = split_article(article.ama_name, ama_lines, codepage)
rendered.update(split_articles) for name, lines in split_articles.items():
return rendered rendered[name] = lines
words = extract_words(lines)
if not words:
continue
word_set = word_index.setdefault(name, set())
word_set.update(words)
inverted_index: Dict[str, set[str]] = {}
for filename, words in word_index.items():
for word in words:
inverted_index.setdefault(word, set()).add(filename)
return rendered, inverted_index
def rewrite_links(markdown: str, base_dir: Path, articles: Dict[Path, Article]) -> str: def rewrite_links(markdown: str, base_dir: Path, articles: Dict[Path, Article]) -> str:
@ -476,13 +621,15 @@ def assemble_files(ama_contents: Dict[str, List[str]], title: str | None, codepa
files.append(("TITLE", title.encode("ascii", "ignore")[:64])) files.append(("TITLE", title.encode("ascii", "ignore")[:64]))
high_bit_used = False high_bit_used = False
articles = dict(ama_contents)
index_bytes = encode_ama("INDEX.AMA", ama_contents.pop("INDEX.AMA"), codepage) index_lines = articles.pop("INDEX.AMA")
index_bytes = encode_ama("INDEX.AMA", index_lines, codepage)
files.append(("INDEX.AMA", index_bytes)) files.append(("INDEX.AMA", index_bytes))
if _has_high_bit(index_bytes): if _has_high_bit(index_bytes):
high_bit_used = True high_bit_used = True
for name, lines in sorted(ama_contents.items()): for name, lines in sorted(articles.items()):
data = encode_ama(name, lines, codepage) data = encode_ama(name, lines, codepage)
files.append((name, data)) files.append((name, data))
if not high_bit_used and _has_high_bit(data): if not high_bit_used and _has_high_bit(data):

Binary file not shown.