indexer
This commit is contained in:
parent
2db4eab94c
commit
ea43375cb4
5 changed files with 415 additions and 8 deletions
|
|
@ -19,6 +19,7 @@ uv run mambler.py --title "Your Book Title" --codepage 437 path/to/index.md outp
|
||||||
- `--title` is optional; when provided the value is embedded in the AMB archive header (truncated to 64 ASCII bytes).
|
- `--title` is optional; when provided the value is embedded in the AMB archive header (truncated to 64 ASCII bytes).
|
||||||
- `--codepage` controls the 8-bit encoding used for every AMA article (default: `437`). Any character that cannot be expressed in the chosen codepage aborts the build with a helpful error so you can pick a better fit.
|
- `--codepage` controls the 8-bit encoding used for every AMA article (default: `437`). Any character that cannot be expressed in the chosen codepage aborts the build with a helpful error so you can pick a better fit.
|
||||||
- If any emitted byte lives in the 0x80–0xFF range, `mambler` automatically writes a companion `UNICODE.MAP` file describing the high-half character mapping, mirroring the recommendation in the AMA/AMB specification.
|
- If any emitted byte lives in the 0x80–0xFF range, `mambler` automatically writes a companion `UNICODE.MAP` file describing the high-half character mapping, mirroring the recommendation in the AMA/AMB specification.
|
||||||
|
- Words of length 2–17 are indexed into `DICT.IDX` so readers can offer fast full-text search. The index is omitted if it would overflow the 64 KiB LoW data limit mandated by the spec.
|
||||||
- The command prints the path of the generated AMB file on success.
|
- The command prints the path of the generated AMB file on success.
|
||||||
|
|
||||||
### Development Notes
|
### Development Notes
|
||||||
|
|
|
||||||
259
format.txt
Normal file
259
format.txt
Normal file
|
|
@ -0,0 +1,259 @@
|
||||||
|
|
||||||
|
==== AMB FORMAT SPECIFICATION ====
|
||||||
|
|
||||||
|
last updated: 2025-08-26
|
||||||
|
|
||||||
|
The latest version of this file can be found on the AMB project's homepage:
|
||||||
|
<http://mateusz.fr/amb/>
|
||||||
|
|
||||||
|
An AMB file (Ancient Machine Book) is an extremely lightweight file format
|
||||||
|
meant to store any kind of hypertext documentation that may be comfortably
|
||||||
|
viewed even on the most ancient PCs: technical manuals, books, etc. Think of
|
||||||
|
it as a retro equivalent of a *.CHM help file. The AMB format is designed to
|
||||||
|
allow for some limited formatting, support internal links and require very
|
||||||
|
little processing power to read, so a reader may be run even on the oldest
|
||||||
|
IBM PC. The format also strives for simplicity of implementation.
|
||||||
|
|
||||||
|
Table Of Contents:
|
||||||
|
|
||||||
|
* The AMB container
|
||||||
|
* Title
|
||||||
|
* AMA format
|
||||||
|
* Codepage encoding
|
||||||
|
* Index data
|
||||||
|
* Rationale
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
|
|
||||||
|
THE AMB CONTAINER
|
||||||
|
|
||||||
|
The AMB file is a container - one could say it is a very simplistic archive
|
||||||
|
format. It starts with a 4-bytes format signature (magic value) "AMB1". Then
|
||||||
|
comes a 2-bytes number that tells how many files are present in the container,
|
||||||
|
followed by the list of all files: each file is described by a file entry.
|
||||||
|
All values are little-endian.
|
||||||
|
|
||||||
|
offset
|
||||||
|
0 format signature: "AMB1"
|
||||||
|
4 files count (16-bit value)
|
||||||
|
6 FILE ENTRY #1
|
||||||
|
FILE ENTRY #2
|
||||||
|
FILE ENTRY #3
|
||||||
|
....
|
||||||
|
DATA
|
||||||
|
|
||||||
|
Each file entry is a 20-bytes structure:
|
||||||
|
|
||||||
|
offset
|
||||||
|
0 filename, 12 characters, zero-padded ("FILE.EXT\0\0\0\0")
|
||||||
|
12 offset where this file starts (32 bits)
|
||||||
|
16 file length, in bytes (16 bits)
|
||||||
|
18 BSD sum (16-bit) of the file
|
||||||
|
|
||||||
|
The AMB archive is expected to contain a set of AMA (Ancient Machine Article)
|
||||||
|
files, and optionally a title file, an index dictionary and a codepage map.
|
||||||
|
AMA files may be compressed with the MVCOMP algorithm, in which case they are
|
||||||
|
named with the "*.AMC" extension.
|
||||||
|
|
||||||
|
An AMB archive must contain at least one article file named either "index.ama"
|
||||||
|
or "index.amc" - this is the first file that an AMB reader will try loading.
|
||||||
|
|
||||||
|
Note: Names of files contained in an AMB archive are to be processed in a case
|
||||||
|
insensitive way and must be composed exclusively of 7-bit characters.
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
|
|
||||||
|
DOCUMENT TITLE
|
||||||
|
|
||||||
|
The AMB title is a string that may be displayed as the document's main title.
|
||||||
|
To set such title, the AMB archive has to contain a file named simply 'title'
|
||||||
|
that would contain the text. The title string should not be longer than
|
||||||
|
64 characters, anything longer might be truncated by the reader.
|
||||||
|
|
||||||
|
The title of the document is expected to be encoded with the same codepage as
|
||||||
|
all the articles. See codepage encoding.
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
|
|
||||||
|
AMA FORMAT
|
||||||
|
|
||||||
|
The AMA format is a text-based file format. For guaranteed interoperability
|
||||||
|
with old machines, its maximum allowed size is 65535 bytes (ie. 2^16 - 1).
|
||||||
|
Larger contents must be segmented into a set of two or more AMA articles.
|
||||||
|
|
||||||
|
An AMB reader must display content with a 78-characters width, hence an AMA
|
||||||
|
article must not contain any line longer than 78 displayable characters. Lines
|
||||||
|
longer than this limit may be truncated by the client reader.
|
||||||
|
|
||||||
|
AMA articles may contain control codes. A control code is a characters pair,
|
||||||
|
where the first is a percent (%) character. Possible control codes:
|
||||||
|
|
||||||
|
%t normal text follows (default state)
|
||||||
|
%h heading follows
|
||||||
|
%l link follows (filename ended by a ':', followed by a description)
|
||||||
|
%! notice/warning follows
|
||||||
|
%b boring text follows (usually displayed grey on grey)
|
||||||
|
%% display a percent character (%)
|
||||||
|
|
||||||
|
It is important to note that the current text mode is reset to %t at the end
|
||||||
|
of every line, hence there is no need to prefix a line of text with %t.
|
||||||
|
|
||||||
|
Line endings may be either LF or CR/LF. The former is recommended, as it is
|
||||||
|
more compact.
|
||||||
|
|
||||||
|
TAB control codes (ASCII decimal value 9) are NOT allowed in AMA files.
|
||||||
|
|
||||||
|
Whenever an external URL appears in an AMA file (for example a link to a web
|
||||||
|
page, to a ftp resource or to a gopher hole) it is encouraged to be enclosed
|
||||||
|
between <> characters. Example: <http://mateusz.fr/amb>. This is a typesetting
|
||||||
|
recommendation based on RFC 3986, it is not part of the AMA specification.
|
||||||
|
Following it would, however, make it much easier for modern AMB readers to
|
||||||
|
detect such links automatically and make them clickable.
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
|
|
||||||
|
CODEPAGE ENCODING
|
||||||
|
|
||||||
|
Since ancient computers are displaying text as 8-bit characters due to the
|
||||||
|
design of early video adapters, AMA files are expected to contain 8-bit text
|
||||||
|
as well. The exact codepage is unspecified by this format definition and
|
||||||
|
depends on the document's target audience.
|
||||||
|
|
||||||
|
To ease displaying of AMB books on modern (unicode-enabled) platforms, any AMB
|
||||||
|
file that contains non-7-bit characters SHOULD also contain a file named
|
||||||
|
"unicode.map". This file contains a sequence of 128 16-bit values, mapping
|
||||||
|
bytes of the range 128..255 into unicode datapoint values. Such file can be
|
||||||
|
readily output by the utf8tocp program <http://mateusz.fr/utf8tocp/>.
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
|
|
||||||
|
INDEX DATA
|
||||||
|
|
||||||
|
On top of AMA files, the AMB archive may contain a file named DICT.IDX. This
|
||||||
|
file, if it exists, provides indexing metadata to allow the client to perform
|
||||||
|
fast and efficient full-text searches across the AMB book.
|
||||||
|
|
||||||
|
The index file contains a hash table: a serie of 256 16-bit indexes, where
|
||||||
|
each index points to a region of the index structure that contains a list of
|
||||||
|
words (LoW). The index (0..255) itself is an 8 bits hash based on the length
|
||||||
|
of the word and its characters. The checksum is made of two nibbles: LC.
|
||||||
|
The high nibble (L) is the length of the word minus 2, while the low nibble
|
||||||
|
(C) is a simple checksum of all the word's characters XORed together. This
|
||||||
|
algorithm can be formalized as follows:
|
||||||
|
|
||||||
|
((wordlen - 2) << 4) | ((a & 15) XOR (b & 15) XOR (...))
|
||||||
|
|
||||||
|
For example, the word "Disk" would end up being indexed under value 0x25,
|
||||||
|
because:
|
||||||
|
|
||||||
|
((4 - 2) << 4) | ((D & 15) XOR (i & 15) XOR (s & 15) XOR (k & 15))
|
||||||
|
translates to: (2 << 4) | (4 XOR 9 XOR 3 XOR 11)
|
||||||
|
which leads to: 32 | 5
|
||||||
|
resulting in: 37 = 0x25
|
||||||
|
|
||||||
|
After the index we can find the pointer to the words list. A pointer is a 16
|
||||||
|
bits file offset from the index structure start.
|
||||||
|
|
||||||
|
It needs to be noted that words of less than 2 characters and more than 17
|
||||||
|
characters cannot be indexed. The presented algorithm has also the interesting
|
||||||
|
side-effect of indexing low and high caps of the ranges a..z and A..Z
|
||||||
|
identically. An important limitation is the fact that the list of words (LoW)
|
||||||
|
is restricted by the 16-bit addressing offset, which means that all LoWs must
|
||||||
|
start at an offset within the first 64 KiB of the file.
|
||||||
|
|
||||||
|
Now that we know the offset at which our LoW starts, we can read the words.
|
||||||
|
First go to the offset, and read a single 16 bits word. Its value contains the
|
||||||
|
number of words in the list. Then, read the words one after another (note that
|
||||||
|
all words in the list have the same length, and you know this length already).
|
||||||
|
Words are always written in lower case characters. Each word is followed by a
|
||||||
|
1-byte value that tells how many files the word has been found in. Then, that
|
||||||
|
many 32-bit file identifiers follow.
|
||||||
|
|
||||||
|
index format:
|
||||||
|
|
||||||
|
* List of words
|
||||||
|
|
||||||
|
xx number of words in the list
|
||||||
|
? word
|
||||||
|
x how many files the word is present in
|
||||||
|
xxxx file identifier 1
|
||||||
|
xxxx file identifier 2
|
||||||
|
...
|
||||||
|
xxxx file identifier n
|
||||||
|
|
||||||
|
(other 255 lists of words follow)
|
||||||
|
|
||||||
|
* hash table
|
||||||
|
|
||||||
|
xx offset of the LoW for words that match hash 0x00
|
||||||
|
xx offset of the LoW for words that match hash 0x01
|
||||||
|
...
|
||||||
|
xx offset of the LoW for words that match hash 0xff
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
|
|
||||||
|
RATIONALE
|
||||||
|
|
||||||
|
The AMB format is, by design, burdened by several limitations. These
|
||||||
|
limitations might be misunderstood as shortcomings, while in essence the AMB
|
||||||
|
format's primary objective is to stay as primitive as possible - so it is easy
|
||||||
|
(and fast) for software to parse and display. Below are listed some of these
|
||||||
|
limitations, with explanations about the reasons that led to them.
|
||||||
|
|
||||||
|
* Line length limited to 78 characters
|
||||||
|
|
||||||
|
The hard-coded limit of 78 chars is meant to ensure that the reader will
|
||||||
|
not have to worry about line wrapping, which highly simplifies the reader's
|
||||||
|
code thus allowing for faster processing and minimizing potential bugs. It
|
||||||
|
is also meant to allow the content creator to design his screens in a
|
||||||
|
deterministic way - that is, without any risk that his semigraphic tables,
|
||||||
|
ASCII drawings or overall screen disposition will be broken by a reader
|
||||||
|
that attempts to rewrap the text at an unpredictable width.
|
||||||
|
The 80-columns width was ubiquitous since the early 80' and seems to be a
|
||||||
|
reasonable baseline expectation, and a 78-characters limit allows the
|
||||||
|
reader to use two columns for its own needs (vertical cursor, border, etc).
|
||||||
|
|
||||||
|
* No control over style (colors) applied to text
|
||||||
|
|
||||||
|
The AMB format defines a set of semantic tags (like "%h" for "heading"). It
|
||||||
|
does not allow control over the exact colors or attributes that will be
|
||||||
|
used by the output device to render the document. This is designed on
|
||||||
|
purpose: AMB documents should be displayable also on monochrome devices.
|
||||||
|
There may also be devices that allow for text attributes like "underlined",
|
||||||
|
"bold", etc - it is up to the AMB reader to make sure the semantic tags are
|
||||||
|
translated into colors/shades/attributes combinations that are nicely
|
||||||
|
rendered on the target hardware.
|
||||||
|
|
||||||
|
* Article size limit of 64 KiB
|
||||||
|
|
||||||
|
A single article (AMA file) is limited to a maximum length of 64 KiB (minus
|
||||||
|
one byte). This limitation makes it easier for MS-DOS readers to load the
|
||||||
|
content: in real-mode Intel memory models, a single memory segment is
|
||||||
|
addressable via 16-bit offsets, hence processing content larger than 64 KiB
|
||||||
|
becomes tricky, as it involves crossing memory segment boundaries, or
|
||||||
|
relying on some kludges like "huge" memory pointers (slow), or dynamically
|
||||||
|
reloading parts of the file from disk (very slow). 64 KiB still allows for
|
||||||
|
more than 30 pages of 80x25 packed text, which should be more than enough
|
||||||
|
even for very complex subjects (and larger contents should simply be
|
||||||
|
dispatched into two or more different articles, which can only be
|
||||||
|
beneficial for readability).
|
||||||
|
|
||||||
|
* Maximum number of 65535 articles
|
||||||
|
|
||||||
|
An AMB book may contain up to 65535 articles and not a single more, because
|
||||||
|
the number of articles is written as a 16-bit integer in the file's header.
|
||||||
|
This allows AMB software to use 16-bit integers when addressing the
|
||||||
|
articles, which is very convenient (and fast) for platforms with 16-bit
|
||||||
|
CPUs. And honestly - is that really a limitation? Even the entire Bible has
|
||||||
|
"only" 1189 chapters, or 31103 verses.
|
||||||
|
|
||||||
|
* Short filenames + low-ascii characters only
|
||||||
|
|
||||||
|
Filenames inside an AMB container are limited to 12 (8+3) characters so
|
||||||
|
an AMB container can be unpacked on an old MS-DOS system.
|
||||||
|
The filenames must contain only low-ascii (7-bit) characters -- for two
|
||||||
|
reasons: so it is possible to unpack an AMB container on any filesystem,
|
||||||
|
independently of the codepage said filesystem relies on, and to make it
|
||||||
|
possible to reliably perform case-insensitive matching of filenames.
|
||||||
|
|
||||||
|
==============================================================================
|
||||||
BIN
long.amb
BIN
long.amb
Binary file not shown.
163
mambler.py
163
mambler.py
|
|
@ -5,10 +5,11 @@ import argparse
|
||||||
import codecs
|
import codecs
|
||||||
import re
|
import re
|
||||||
import struct
|
import struct
|
||||||
|
import sys
|
||||||
from collections import deque
|
from collections import deque
|
||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Callable, Dict, Iterable, List, Tuple
|
from typing import Callable, Dict, Iterable, List, Optional, Tuple
|
||||||
|
|
||||||
from md2txt import convert_markdown
|
from md2txt import convert_markdown
|
||||||
from md2txt.conversion.core import parse_frontmatter
|
from md2txt.conversion.core import parse_frontmatter
|
||||||
|
|
@ -249,6 +250,114 @@ def _has_high_bit(data: bytes) -> bool:
|
||||||
return any(byte >= 0x80 for byte in data)
|
return any(byte >= 0x80 for byte in data)
|
||||||
|
|
||||||
|
|
||||||
|
def compute_file_offsets(files: List[Tuple[str, bytes]], include_dict: bool) -> Dict[str, int]:
|
||||||
|
total_files = len(files) + (1 if include_dict else 0)
|
||||||
|
offset = 6 + 20 * total_files
|
||||||
|
mapping: Dict[str, int] = {}
|
||||||
|
for name, data in files:
|
||||||
|
mapping[name.upper()] = offset
|
||||||
|
offset += len(data)
|
||||||
|
return mapping
|
||||||
|
|
||||||
|
|
||||||
|
def build_dict_index(
|
||||||
|
word_index: Dict[str, set[str]],
|
||||||
|
file_offsets: Dict[str, int],
|
||||||
|
codepage: CodepageInfo,
|
||||||
|
) -> Optional[bytes]:
|
||||||
|
bucket_words: Dict[int, Dict[str, Tuple[bytes, List[int]]]] = {}
|
||||||
|
for word, filenames in word_index.items():
|
||||||
|
try:
|
||||||
|
encoded_word = codepage.encode(word)
|
||||||
|
except UnicodeEncodeError:
|
||||||
|
continue
|
||||||
|
length = len(encoded_word)
|
||||||
|
if not (2 <= length <= 17):
|
||||||
|
continue
|
||||||
|
bucket = compute_word_hash(encoded_word)
|
||||||
|
file_ids = sorted({file_offsets[name.upper()] for name in filenames if name.upper() in file_offsets})
|
||||||
|
if not file_ids:
|
||||||
|
continue
|
||||||
|
bucket_map = bucket_words.setdefault(bucket, {})
|
||||||
|
bucket_map[word] = (encoded_word, file_ids)
|
||||||
|
|
||||||
|
if not bucket_words:
|
||||||
|
return None
|
||||||
|
|
||||||
|
lows = bytearray()
|
||||||
|
offsets: List[int] = []
|
||||||
|
|
||||||
|
for bucket in range(256):
|
||||||
|
offsets.append(len(lows))
|
||||||
|
entries = bucket_words.get(bucket)
|
||||||
|
if not entries:
|
||||||
|
lows.extend(struct.pack("<H", 0))
|
||||||
|
continue
|
||||||
|
|
||||||
|
sorted_entries = sorted(entries.items(), key=lambda item: item[0])
|
||||||
|
word_length = len(sorted_entries[0][1][0])
|
||||||
|
lows.extend(struct.pack("<H", len(sorted_entries)))
|
||||||
|
for word, (encoded_word, file_ids) in sorted_entries:
|
||||||
|
if len(encoded_word) != word_length:
|
||||||
|
raise ValueError(f"Inconsistent word length for hash bucket 0x{bucket:02x}.")
|
||||||
|
if len(file_ids) > 255:
|
||||||
|
raise ValueError(f"Word '{word}' appears in more than 255 files, cannot encode index.")
|
||||||
|
lows.extend(encoded_word)
|
||||||
|
lows.append(len(file_ids))
|
||||||
|
for file_id in file_ids:
|
||||||
|
lows.extend(struct.pack("<I", file_id))
|
||||||
|
|
||||||
|
if len(lows) >= 0x10000:
|
||||||
|
raise ValueError("Generated DICT.IDX exceeds 64 KiB limit for word lists.")
|
||||||
|
|
||||||
|
hash_table = bytearray()
|
||||||
|
for offset in offsets:
|
||||||
|
hash_table.extend(struct.pack("<H", offset))
|
||||||
|
|
||||||
|
return bytes(lows + hash_table)
|
||||||
|
|
||||||
|
|
||||||
|
def compute_word_hash(encoded_word: bytes) -> int:
|
||||||
|
length = len(encoded_word)
|
||||||
|
checksum = 0
|
||||||
|
for byte in encoded_word:
|
||||||
|
checksum ^= (byte & 0x0F)
|
||||||
|
return ((length - 2) << 4) | (checksum & 0x0F)
|
||||||
|
|
||||||
|
|
||||||
|
WORD_MIN_LENGTH = 2
|
||||||
|
WORD_MAX_LENGTH = 17
|
||||||
|
|
||||||
|
|
||||||
|
def extract_words(lines: List[str]) -> set[str]:
|
||||||
|
words: set[str] = set()
|
||||||
|
for line in lines:
|
||||||
|
stripped = _strip_control_codes(line)
|
||||||
|
buffer: List[str] = []
|
||||||
|
for char in stripped:
|
||||||
|
if char.isalnum():
|
||||||
|
buffer.append(char.lower())
|
||||||
|
continue
|
||||||
|
if len(buffer) >= WORD_MIN_LENGTH:
|
||||||
|
word = "".join(buffer)
|
||||||
|
if WORD_MIN_LENGTH <= len(word) <= WORD_MAX_LENGTH:
|
||||||
|
words.add(word)
|
||||||
|
buffer = []
|
||||||
|
if len(buffer) >= WORD_MIN_LENGTH:
|
||||||
|
word = "".join(buffer)
|
||||||
|
if WORD_MIN_LENGTH <= len(word) <= WORD_MAX_LENGTH:
|
||||||
|
words.add(word)
|
||||||
|
return words
|
||||||
|
|
||||||
|
|
||||||
|
def _strip_control_codes(line: str) -> str:
|
||||||
|
result = line
|
||||||
|
result = re.sub(r"%l[^:]+:", "", result)
|
||||||
|
result = result.replace("%t", "").replace("%!", "").replace("%b", "").replace("%h", "")
|
||||||
|
result = result.replace("%%", "%")
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
@dataclass
|
@dataclass
|
||||||
class Article:
|
class Article:
|
||||||
source: Path
|
source: Path
|
||||||
|
|
@ -287,8 +396,31 @@ def main(argv: Iterable[str] | None = None) -> int:
|
||||||
|
|
||||||
def build_amb(root_markdown: Path, title: str | None, codepage: CodepageInfo) -> bytes:
|
def build_amb(root_markdown: Path, title: str | None, codepage: CodepageInfo) -> bytes:
|
||||||
articles = collect_articles(root_markdown)
|
articles = collect_articles(root_markdown)
|
||||||
ama_contents = render_articles(articles, codepage)
|
ama_contents, word_index = render_articles(articles, codepage)
|
||||||
files = assemble_files(ama_contents, title, codepage)
|
base_files = assemble_files(ama_contents, title, codepage)
|
||||||
|
|
||||||
|
file_offsets = compute_file_offsets(base_files, include_dict=False)
|
||||||
|
try:
|
||||||
|
dict_bytes = build_dict_index(word_index, file_offsets, codepage)
|
||||||
|
except ValueError as exc:
|
||||||
|
print(f"[mambler] Skipping dictionary index: {exc}", file=sys.stderr)
|
||||||
|
dict_bytes = None
|
||||||
|
|
||||||
|
if dict_bytes is None:
|
||||||
|
files = base_files
|
||||||
|
else:
|
||||||
|
adjusted_offsets = compute_file_offsets(base_files, include_dict=True)
|
||||||
|
try:
|
||||||
|
dict_bytes_adjusted = build_dict_index(word_index, adjusted_offsets, codepage)
|
||||||
|
except ValueError as exc:
|
||||||
|
print(f"[mambler] Skipping dictionary index: {exc}", file=sys.stderr)
|
||||||
|
files = base_files
|
||||||
|
else:
|
||||||
|
if dict_bytes_adjusted is None:
|
||||||
|
files = base_files
|
||||||
|
else:
|
||||||
|
files = base_files + [("DICT.IDX", dict_bytes_adjusted)]
|
||||||
|
|
||||||
return pack_amb(files)
|
return pack_amb(files)
|
||||||
|
|
||||||
|
|
||||||
|
|
@ -351,8 +483,9 @@ def assign_ama_name(stem: str, existing: set[str]) -> str:
|
||||||
return name
|
return name
|
||||||
|
|
||||||
|
|
||||||
def render_articles(articles: Dict[Path, Article], codepage: CodepageInfo) -> Dict[str, List[str]]:
|
def render_articles(articles: Dict[Path, Article], codepage: CodepageInfo) -> Tuple[Dict[str, List[str]], Dict[str, set[str]]]:
|
||||||
rendered: Dict[str, List[str]] = {}
|
rendered: Dict[str, List[str]] = {}
|
||||||
|
word_index: Dict[str, set[str]] = {}
|
||||||
|
|
||||||
for path, article in articles.items():
|
for path, article in articles.items():
|
||||||
content = path.read_text(encoding="utf-8")
|
content = path.read_text(encoding="utf-8")
|
||||||
|
|
@ -366,8 +499,20 @@ def render_articles(articles: Dict[Path, Article], codepage: CodepageInfo) -> Di
|
||||||
renderer_name="ama",
|
renderer_name="ama",
|
||||||
)
|
)
|
||||||
split_articles = split_article(article.ama_name, ama_lines, codepage)
|
split_articles = split_article(article.ama_name, ama_lines, codepage)
|
||||||
rendered.update(split_articles)
|
for name, lines in split_articles.items():
|
||||||
return rendered
|
rendered[name] = lines
|
||||||
|
words = extract_words(lines)
|
||||||
|
if not words:
|
||||||
|
continue
|
||||||
|
word_set = word_index.setdefault(name, set())
|
||||||
|
word_set.update(words)
|
||||||
|
|
||||||
|
inverted_index: Dict[str, set[str]] = {}
|
||||||
|
for filename, words in word_index.items():
|
||||||
|
for word in words:
|
||||||
|
inverted_index.setdefault(word, set()).add(filename)
|
||||||
|
|
||||||
|
return rendered, inverted_index
|
||||||
|
|
||||||
|
|
||||||
def rewrite_links(markdown: str, base_dir: Path, articles: Dict[Path, Article]) -> str:
|
def rewrite_links(markdown: str, base_dir: Path, articles: Dict[Path, Article]) -> str:
|
||||||
|
|
@ -476,13 +621,15 @@ def assemble_files(ama_contents: Dict[str, List[str]], title: str | None, codepa
|
||||||
files.append(("TITLE", title.encode("ascii", "ignore")[:64]))
|
files.append(("TITLE", title.encode("ascii", "ignore")[:64]))
|
||||||
|
|
||||||
high_bit_used = False
|
high_bit_used = False
|
||||||
|
articles = dict(ama_contents)
|
||||||
|
|
||||||
index_bytes = encode_ama("INDEX.AMA", ama_contents.pop("INDEX.AMA"), codepage)
|
index_lines = articles.pop("INDEX.AMA")
|
||||||
|
index_bytes = encode_ama("INDEX.AMA", index_lines, codepage)
|
||||||
files.append(("INDEX.AMA", index_bytes))
|
files.append(("INDEX.AMA", index_bytes))
|
||||||
if _has_high_bit(index_bytes):
|
if _has_high_bit(index_bytes):
|
||||||
high_bit_used = True
|
high_bit_used = True
|
||||||
|
|
||||||
for name, lines in sorted(ama_contents.items()):
|
for name, lines in sorted(articles.items()):
|
||||||
data = encode_ama(name, lines, codepage)
|
data = encode_ama(name, lines, codepage)
|
||||||
files.append((name, data))
|
files.append((name, data))
|
||||||
if not high_bit_used and _has_high_bit(data):
|
if not high_bit_used and _has_high_bit(data):
|
||||||
|
|
|
||||||
BIN
output.amb
BIN
output.amb
Binary file not shown.
Loading…
Add table
Add a link
Reference in a new issue