Oregami
Repositories/oxedyne/fe2o3

oxedyne/fe2o3/fe2o3_text/src/doc/html/mod.rs

3.2 KiB, 7 runs

created by r1870400018:14322, which is this file's identity for as long as the history lasts, whatever it is later renamed to

download · who wrote it · its history

1//! HTML -- a reader for the form a generator exports, and a writer for the form a browser reads.
2//!
3//! This module faces both ways. [`parse`] reads HTML into the neutral [document tree](crate::doc), as
4//! [`markdown`](crate::doc::markdown) reads Markdown into it; [`render`] walks the tree back out. The
5//! two are not inverses and are not meant to be -- what the tree does not carry, no reader can invent
6//! and no writer can restore -- but between them they make the tree's neutrality testable rather than
7//! merely asserted.
8//!
9//! # Why read HTML at all
10//!
11//! Because a great deal of prose is *exported* to it rather than written in it. A typesetter that
12//! evaluates an author's own macros and emits HTML has already done the hard half of the work: what
13//! comes out is the prose, with the author's abbreviations, cross-references and templates resolved.
14//! Reading that HTML is how such prose reaches the tree without the reader having to understand the
15//! language it was written in.
16//!
17//! That is what this reads: HTML a generator wrote. It is not a browser's parser, does not recover
18//! from mis-nesting the way a browser must, and does not try to. See [`read`] for what it does with
19//! the tags it does not know, which is the part worth knowing.
20//!
21//! # HTML collapses whitespace
22//!
23//! A run of spaces, tabs and newlines between two words says one space, and the whitespace at either
24//! end of a block says nothing at all. So an exporter's indentation and line endings are not the
25//! author's, and are not kept.
26//!
27//! This is the mirror of the rule [`markdown`](crate::doc::markdown) states, and it exists for the
28//! same reason. There a newline within a paragraph had to *become* a space; here a run of whitespace
29//! has to *collapse to* one. Both say that where a line ended in the source is not something the
30//! author asked for, and that prose should reflow to the width it is read at rather than freeze at the
31//! width it was written or exported at. Get this wrong and every paragraph of a book carries the
32//! exporter's line breaks for ever.
33//!
34//! The one exception is `<pre>`, where whitespace is exactly what the content means. Its line
35//! structure reaches [`Block::Code`](crate::doc::Block::Code) intact.
36//!
37//! [`Inline::Break`](crate::doc::Inline::Break) therefore only ever comes from a `<br>`, and never
38//! from a newline in the source.
39//!
40//! # Usage
41//!
42//! ```ignore
43//! use oxedyne_fe2o3_text::doc::html;
44//!
45//! let tree = res!(html::parse("<h1>A heading</h1>\n<p>A paragraph with <em>emphasis</em>.</p>\n"));
46//! let out = html::render(&tree);
47//! ```
48
49pub mod read;
50pub mod write;
51
52use crate::doc::Doc;
53
54use oxedyne_fe2o3_core::prelude::*;
55
56pub use self::write::{
57 Opts,
58 escape_attr,
59 escape_text,
60 render,
61 render_with,
62};
63
64/// Reads HTML and produces its document tree.
65///
66/// Parsing does not fail on markup that means less than it might: a tag the tree has no node for is
67/// unwrapped, a stray close tag closes nothing, and an element left open runs to the end of what
68/// encloses it. The outcome is an error only when the input breaks the one limit the reader holds
69/// against a hostile document, [`read::DEPTH_LIMIT`].
70pub fn parse(src: &str) -> Outcome<Doc> {
71 let blocks = res!(read::parse(src));
72 Ok(Doc { blocks })
73}