oxedyne/fe2o3/fe2o3_text/src/doc/markdown/mod.rs
2.7 KiB, 10 runs
created by r1870400018:14279, which is this file's identity for as long as the history lasts, whatever it is later renamed to
download · who wrote it · its history
| 1 | //! Markdown -- a reader for the lightweight markup language, producing the neutral |
| 2 | //! [document tree](crate::doc). |
| 3 | //! |
| 4 | //! Markdown is the form most prose is written in, and a great deal of prose already exists in it. This |
| 5 | //! is the front-end that reads it. What it produces belongs to [`crate::doc`] and knows nothing of |
| 6 | //! Markdown, which is what lets a second front-end produce the same tree. |
| 7 | //! |
| 8 | //! # The dialect |
| 9 | //! |
| 10 | //! The subset parsed is the CommonMark core that prose actually uses: ATX and setext headings, |
| 11 | //! paragraphs, fenced and indented code, block quotations, ordered and unordered lists (nested), |
| 12 | //! thematic breaks, GitHub's pipe tables, and the inline run of emphasis, strong emphasis, links, |
| 13 | //! images, code spans and hard breaks. It is not a conformant CommonMark implementation and does not |
| 14 | //! try to be: the reference |
| 15 | //! test suite is largely a catalogue of pathological nesting that no author writes. Where this reader |
| 16 | //! and CommonMark differ on such input, this reader is simply making its own choice. |
| 17 | //! |
| 18 | //! # A soft line break says a space |
| 19 | //! |
| 20 | //! A single newline within a paragraph is a soft break, and it contributes a space rather than a |
| 21 | //! newline. Where an author's editor wrapped a line is not where the author asked for a break, and |
| 22 | //! preserving it would freeze prose at the width it was typed at instead of reflowing to the width it |
| 23 | //! is read at. |
| 24 | //! |
| 25 | //! The trap is that CommonMark's own HTML output writes a soft break as a newline, which looks like a |
| 26 | //! licence to keep it. It is not: HTML collapses whitespace, so that newline is rendered as a space by |
| 27 | //! the reader that receives it. A renderer that honours whitespace would take it as a break the author |
| 28 | //! never asked for. So the tree carries the meaning and not the byte, and |
| 29 | //! [`Inline::Break`](crate::doc::Inline::Break) is only ever a break the author did ask for. |
| 30 | //! |
| 31 | //! # Usage |
| 32 | //! |
| 33 | //! ```ignore |
| 34 | //! use oxedyne_fe2o3_text::doc::markdown; |
| 35 | //! |
| 36 | //! let tree = res!(markdown::parse("# A heading\n\nA paragraph with *emphasis*.\n")); |
| 37 | //! ``` |
| 38 | |
| 39 | pub mod block; |
| 40 | pub mod inline; |
| 41 | pub mod write; |
| 42 | |
| 43 | use crate::doc::Doc; |
| 44 | |
| 45 | use oxedyne_fe2o3_core::prelude::*; |
| 46 | |
| 47 | /// Reads Markdown text and produces its document tree. |
| 48 | /// |
| 49 | /// Parsing does not fail on badly formed markup: Markdown has no syntax errors, only text that means |
| 50 | /// less than the author hoped. An unclosed fence runs to the end, an unmatched bracket is literal |
| 51 | /// text, and a stray asterisk is an asterisk. The outcome is an error only when the input breaks a |
| 52 | /// limit the reader holds against a hostile document, such as nesting past [`block::DEPTH_LIMIT`]. |
| 53 | pub fn parse(src: &str) -> Outcome<Doc> { |
| 54 | let blocks = res!(block::parse(src)); |
| 55 | Ok(Doc { blocks }) |
| 56 | } |