oxedyne/fe2o3/fe2o3_text/src/doc/html/mod.rs
3.2 KiB, 7 runs
created by r1870400018:14322, which is this file's identity for as long as the history lasts, whatever it is later renamed to
download · who wrote it · its history
| 1 | //! HTML -- a reader for the form a generator exports, and a writer for the form a browser reads. |
| 2 | //! |
| 3 | //! This module faces both ways. [`parse`] reads HTML into the neutral [document tree](crate::doc), as |
| 4 | //! [`markdown`](crate::doc::markdown) reads Markdown into it; [`render`] walks the tree back out. The |
| 5 | //! two are not inverses and are not meant to be -- what the tree does not carry, no reader can invent |
| 6 | //! and no writer can restore -- but between them they make the tree's neutrality testable rather than |
| 7 | //! merely asserted. |
| 8 | //! |
| 9 | //! # Why read HTML at all |
| 10 | //! |
| 11 | //! Because a great deal of prose is *exported* to it rather than written in it. A typesetter that |
| 12 | //! evaluates an author's own macros and emits HTML has already done the hard half of the work: what |
| 13 | //! comes out is the prose, with the author's abbreviations, cross-references and templates resolved. |
| 14 | //! Reading that HTML is how such prose reaches the tree without the reader having to understand the |
| 15 | //! language it was written in. |
| 16 | //! |
| 17 | //! That is what this reads: HTML a generator wrote. It is not a browser's parser, does not recover |
| 18 | //! from mis-nesting the way a browser must, and does not try to. See [`read`] for what it does with |
| 19 | //! the tags it does not know, which is the part worth knowing. |
| 20 | //! |
| 21 | //! # HTML collapses whitespace |
| 22 | //! |
| 23 | //! A run of spaces, tabs and newlines between two words says one space, and the whitespace at either |
| 24 | //! end of a block says nothing at all. So an exporter's indentation and line endings are not the |
| 25 | //! author's, and are not kept. |
| 26 | //! |
| 27 | //! This is the mirror of the rule [`markdown`](crate::doc::markdown) states, and it exists for the |
| 28 | //! same reason. There a newline within a paragraph had to *become* a space; here a run of whitespace |
| 29 | //! has to *collapse to* one. Both say that where a line ended in the source is not something the |
| 30 | //! author asked for, and that prose should reflow to the width it is read at rather than freeze at the |
| 31 | //! width it was written or exported at. Get this wrong and every paragraph of a book carries the |
| 32 | //! exporter's line breaks for ever. |
| 33 | //! |
| 34 | //! The one exception is `<pre>`, where whitespace is exactly what the content means. Its line |
| 35 | //! structure reaches [`Block::Code`](crate::doc::Block::Code) intact. |
| 36 | //! |
| 37 | //! [`Inline::Break`](crate::doc::Inline::Break) therefore only ever comes from a `<br>`, and never |
| 38 | //! from a newline in the source. |
| 39 | //! |
| 40 | //! # Usage |
| 41 | //! |
| 42 | //! ```ignore |
| 43 | //! use oxedyne_fe2o3_text::doc::html; |
| 44 | //! |
| 45 | //! let tree = res!(html::parse("<h1>A heading</h1>\n<p>A paragraph with <em>emphasis</em>.</p>\n")); |
| 46 | //! let out = html::render(&tree); |
| 47 | //! ``` |
| 48 | |
| 49 | pub mod read; |
| 50 | pub mod write; |
| 51 | |
| 52 | use crate::doc::Doc; |
| 53 | |
| 54 | use oxedyne_fe2o3_core::prelude::*; |
| 55 | |
| 56 | pub use self::write::{ |
| 57 | Opts, |
| 58 | escape_attr, |
| 59 | escape_text, |
| 60 | render, |
| 61 | render_with, |
| 62 | }; |
| 63 | |
| 64 | /// Reads HTML and produces its document tree. |
| 65 | /// |
| 66 | /// Parsing does not fail on markup that means less than it might: a tag the tree has no node for is |
| 67 | /// unwrapped, a stray close tag closes nothing, and an element left open runs to the end of what |
| 68 | /// encloses it. The outcome is an error only when the input breaks the one limit the reader holds |
| 69 | /// against a hostile document, [`read::DEPTH_LIMIT`]. |
| 70 | pub fn parse(src: &str) -> Outcome<Doc> { |
| 71 | let blocks = res!(read::parse(src)); |
| 72 | Ok(Doc { blocks }) |
| 73 | } |