--- title: "The Saive markdown schema" description: "The YAML frontmatter and body conventions in every file Saive writes. Published so you can read, validate, and rebuild your library without us." version: "1" updatedAt: "2026-08-14" slug: "v1" --- Every save in Saive is a markdown file with YAML frontmatter. Not a row in a database that can be exported to markdown, but a markdown file, written first, that a database happens to index. This page is the specification for those files. It exists so that you can write your own tooling against your library, verify that an export is complete, or walk away entirely and still have something useful. It describes what the code actually emits today, not what's planned. Where the two differ, there's a [known gaps](#known-gaps) section at the bottom that says so plainly. ## The contract 1. **The file is authoritative.** Every piece of per-save state Saive relies on round-trips through frontmatter and body. The database is an index, and rebuilding that index from the vault files is a real operation in the codebase, not a design aspiration. What a rebuild does not restore is listed in [what doesn't travel](#what-doesnt-travel). 2. **Identifiers are stable.** A save's `uuid` is assigned at creation and never changes, not on retitle, retag, move, or folder rename. It is also the filename. 3. **Standard keys where standards exist.** Obsidian's reserved keys (`tags`, `aliases`), the Pandoc / Dublin Core overlap (`title`, `author`, `description`, `source`), and the common bookmarking keys (`url`, `date_saved`, `status`) are used as-is. 4. **Extensions are namespaced.** Anything Saive-specific is prefixed `saive_`. Every one of those keys is safe to ignore, and safe to delete. 5. **Nothing in a file points at Saive.** No API endpoints, no callback URLs, no identifiers that only resolve against our servers. The two current exceptions are named in [what doesn't travel](#what-doesnt-travel). ## File layout A save is a single `.md` file named for its UUID: ```text Recipes/3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d.md 7e8f9a1b-2c3d-4e5f-6789-0abcdef12345.md _concepts/edge-computing-c9d8.md throughlines/0042-2026-07-04.md ``` Folders are real directories. A save with no folder sits at the root. Concept files live under `_concepts/`, throughlines under `throughlines/`; both names are reserved and can't be used as folder names. Naming files by UUID rather than by title is deliberate: it means a bare `[[3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d]]` wikilink resolves in Obsidian no matter which folder the file has been moved to, and a retitle is never a rename. Human-readable titles live in `title` and `aliases`, which is where Obsidian's quick switcher looks anyway. ## A complete file ```markdown --- tags: - edge-compute - infrastructure aliases: - The Quiet Revolution in Edge title: The Quiet Revolution in Edge Compute author: Jane Doe description: >- Edge computing has shifted from CDN supplement to a foundational layer of cloud infrastructure. source: example.com url: https://example.com/2026/05/edge-compute date_saved: "2026-05-17T14:32:00.000Z" date_published: "2026-05-15T08:00:00.000Z" status: unread reading_time_min: 8 uuid: 3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d created: "2026-05-17T14:32:00.000Z" modified: "2026-05-17T14:33:12.000Z" saive_schema: 1 saive_kind: article saive_visibility: private saive_summary: >- The article traces three inflection points in edge computing and argues that the edge is no longer a specialised tier but the centre of gravity for new infrastructure. saive_concepts: - "[[Edge Computing]]" saive_starred: true saive_source_metadata: og_image_url: https://example.com/img/edge.jpg favicon_url: https://example.com/favicon.ico saive_extraction: tier: full source: extension paywall_detected: false body_present: true captured_at: "2026-05-17T14:32:00.000Z" saive_processing: pipeline_version: 1 ai_processed: true skipped_reason: null summary_model: claude-sonnet-4-6 tag_model: claude-haiku-4-5 processed_at: "2026-05-17T14:33:12.000Z" --- # The Quiet Revolution in Edge Compute For most of the last decade, "edge computing" meant a CDN with a few lightweight execution primitives bolted on. ==That framing is now obsolete==[^hl-1] ## Notes Worth revisiting when we plan next year's infra budget. [^hl-1]: Compare against the 2024 latency numbers. ``` ## Field reference Fields marked optional are **omitted entirely** when absent rather than written as `null` or an empty array. ### Obsidian-reserved | Field | Required | Type | Notes | |---|---|---|---| | `tags` | yes | string[] | Lowercase, hyphenated. May be empty. | | `aliases` | no | string[] | Alternative titles. When Saive rewrites an unusable captured title, the original is appended here. | ### Standard interop | Field | Required | Type | Notes | |---|---|---|---| | `title` | yes | string | | | `author` | no | string | From the source page or AI extraction. | | `description` | no | string | The publisher's own one-or-two-line dek. Distinct from `saive_summary`. | | `source` | yes | string | Currently the domain. See [known gaps](#known-gaps). | ### Bookmark fields | Field | Required | Type | Notes | |---|---|---|---| | `url` | yes | URL | The canonical URL of the saved page. | | `date_saved` | yes | ISO-8601 | When you saved it. | | `date_published` | no | string | The publisher's timestamp, stored verbatim as a string, not reformatted. | | `status` | yes | `unread` · `read` · `archived` | | | `reading_time_min` | no | integer | Estimate. | ### Identity | Field | Required | Type | Notes | |---|---|---|---| | `uuid` | yes | UUIDv4 | Assigned at creation, never changes. Matches the filename. | | `created` | yes | ISO-8601 | Never changes. | | `modified` | yes | ISO-8601 | Updated on every write. | ### Saive extensions | Field | Required | Type | Notes | |---|---|---|---| | `saive_schema` | yes | integer | Currently `1`. | | `saive_kind` | yes | `article` · `concept` · `throughline` | Which of the three file shapes this is. | | `saive_visibility` | yes | `private` · `public` | Whether the save appears on a public folder feed. Absent is read as `private`. | | `saive_shared_at` | no | ISO-8601 datetime | When the save last turned public. Present only while `saive_visibility` is `public`. | | `saive_summary` | no | string | The AI summary: what the piece argues. | | `saive_summary_descriptive` | no | string | A neutral description of what the piece *is*. Used on shared surfaces where the substantive summary stays private to whoever saved it. | | `saive_concepts` | no | string[] | Obsidian wikilinks, e.g. `"[[Edge Computing]]"`. These are what produce real edges in the graph view. | | `saive_starred` | no | `true` | Present only when starred. | | `saive_user_edited_summary` | no | `true` | Present only when you've edited the summary yourself, so a rebuild doesn't overwrite your words with a regenerated one. | | `saive_source_metadata` | no | object | `og_image_url`, `og_image_caption`, `favicon_url`. Emitted only when at least one is known. | | `saive_extraction` | yes | object | How the content was obtained. Below. | | `saive_processing` | yes | object | Whether AI ran, and what ran. Below. | | `saive_archetype` | no | object | What the content *is*. `type` holds the slug (`recipe`, `video`, …), with that archetype's data fields flat beneath it. Keys Saive doesn't recognise are preserved on read rather than dropped. | | `saive_source` | no | object | `id` + `parser_version` for the origin-specific parser that handled the save. | | `saive_import` | no | object | Present only on bulk-imported saves. `format` + `session_id` + `imported_at`. Below. | ### saive_import Present only on saves that arrived through bulk import, and absent on everything you capture with the extension, the web app, or the MCP server. | Field | Required | Type | Notes | |---|---|---|---| | `format` | yes | `netscape` · `instapaper` · `pocket` · `raindrop` · `readwise` | The export the save came from. | | `session_id` | yes | UUID | The import run. It means nothing outside Saive; safe to delete, safe to ignore. | | `imported_at` | yes | ISO-8601 | When Saive ingested the file. | **The timestamp rule is the part worth knowing.** On an imported save, `created` and `date_saved` carry the *source's* saved-at date — a link you saved in Pocket in 2014 still reads as 2014, because the history is most of what a library is worth. `imported_at` is the only field that records when the file actually arrived. It's also how Saive keeps a five-thousand-item import from burying everything you saved last week. This block is orthogonal to `saive_source` below, which records which parser extracted the *content*: an imported YouTube link carries both. ### saive_extraction Records how the body in this file was obtained, so a thin save is distinguishable from a real capture without guessing. | Field | Required | Type | Notes | |---|---|---|---| | `tier` | yes | `full` · `preview` · `metadata_only` · `url_only` | Quality of the capture. | | `source` | yes | `extension` · `server` | Browser extension capture, or server-side fetch. | | `paywall_detected` | yes | boolean | | | `body_present` | yes | boolean | Whether the article body is actually in this file. Must agree with the file; a mismatch is a corruption signal. | | `captured_at` | yes | ISO-8601 | | | `format` | no | `pdf` | Present when the body was converted from another document format. Its absence means ordinary HTML extraction. | ### saive_processing | Field | Required | Type | Notes | |---|---|---|---| | `pipeline_version` | yes | integer | | | `ai_processed` | yes | boolean | Whether AI ran and produced metadata for this save. | | `skipped_reason` | yes | `over_quota` · `absent` · `failed` · `null` | `null` when `ai_processed` is true. | | `summary_model` | when processed | string | | | `tag_model` | when processed | string | | | `processed_at` | when processed | ISO-8601 | | | `title_model` | no | string | Set when AI replaced an unusable captured title. Independent of `ai_processed`, so a link-only save can get a title fix without any summary. | | `title_source` | no | `pdf_heading` · `pdf_filename` | Low-confidence provenance for a heuristically derived PDF title. | | `embedding_ref` | no | string | A pointer into Saive's search index. See [what doesn't travel](#what-doesnt-travel). | | `archetype_confidence` | no | number 0–1 | | | `archetype_detected_by` | no | `schema_org` · `og` · `url_pattern` · `llm` · `user` | `user` is never overwritten by automated detection. | ## Body conventions The body is the extracted article, converted to markdown with its structure intact: headings, lists, code blocks, blockquotes, tables. **H1.** The first line is an H1 matching `title`. **Highlights** are Obsidian-flavoured `==highlight==` markers inline in the prose. A highlight that carries a note gets a footnote reference immediately after the closing `==`, in the form `[^hl-1]`, with its definition at the end of the file. Publisher footnotes (`[^1]`, `[^note]`) are left alone; only the `hl-` prefix is ours. Highlights are not stored in frontmatter; the marked-up body *is* the storage. **Your notes** live under a heading with an explicit sentinel: ```markdown ## Notes ``` The sentinel is the only rule. There's no heuristic. A `## Notes` heading without it is treated as article content, so an article that happens to have a "Notes" section is never mistaken for your annotation. Everything after the sentinel is yours; Saive's AI never rewrites it. ## Concept files Concepts are clusters Saive derives across your saves, one file each under `_concepts/`, with `saive_kind: concept`. Each has two marked regions: ```markdown ## Contributing saves - [[3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d]] ## Your notes Anything you write here survives regeneration. ``` Only the text between the auto-managed sentinels is ever rewritten. Content outside it, including but not limited to the user block, is preserved verbatim. If the auto-managed sentinels are missing from a file, Saive refuses to guess where the list belongs rather than overwriting your edits. The wikilinks are what make this useful outside Saive: drop the folder into Obsidian and the graph view shows articles connected to the concepts they contributed to, with no plugin and no import step. ## Throughlines Periodic AI syntheses of recent saves, one file per issue under `throughlines/`, with `saive_kind: throughline`. Alongside the identity fields they carry `issue`, `published`, `coverage_start` / `coverage_end`, `themes`, `generated_by`, and a `sources` list mapping each paragraph of the narrative back to the UUIDs it drew on. Always written `saive_visibility: private`. ## Reading rules If you're writing something that consumes these files: **Parse leniently.** Saive's own reader does. Unknown keys, `saive_`-prefixed or not, are preserved, not dropped. An unrecognised value in a constrained field degrades to "unset" rather than throwing. **Identity is the only hard requirement.** `uuid`, `title`, `url`, `source`, `created`, `status`, and `tags` must be present and well-formed. Everything else has a defined fallback. **Accept two shapes for every timestamp.** Saive writes them quoted, so YAML reads them back as strings. Obsidian's property editor rewrites frontmatter and strips the quotes, after which the same values parse as YAML timestamps: `Date` objects under js-yaml and PyYAML, not strings. This is expected, not corruption. Handle both. Along the same lines: don't assume a file that has passed through another editor is byte-identical to what Saive wrote. `null` becomes blank, folded scalars reflow. **Wikilinks may lose their brackets.** `saive_concepts` entries are written as `"[[Label]]"`, but a hand-edited file may carry a bare `Label`. Accept both. ## What doesn't travel Some things are honestly not portable, and it's better to name them than to let you find out: **`saive_processing.embedding_ref`** is a pointer into Saive's vector index. It means nothing outside Saive and recomputing it requires re-embedding the text. It is safe to delete and safe to ignore. **`saive_source.id`** names a parser: the code that handled a YouTube page or an arXiv paper differently from a generic article. Archetype *definitions* are declarative data and travel fine, but parsers are executable code and don't serialise into a markdown file. Outside Saive that field is an opaque string. **Sharing state isn't in the files at all.** Folder sharing levels, public URLs, feed tokens and invited-reader grants live only in the index. If the index is rebuilt from vault files, folders come back private and member-less. That fails closed by design: the worst case is a dead share link, never an accidental exposure. Empty folders disappear in the same way, since a folder in the vault is only ever a directory with saves in it. Everything else in the file is either self-describing or standard. ## Versioning `saive_schema` is an integer, currently `1`. Additive optional fields don't bump it; the readers that matter, including any you write, should ignore fields they don't know. It bumps only when the meaning of an existing field changes, and that comes with a migration. A field's meaning is never silently repurposed. The revision of *this document* is tracked separately, in the changelog below, so an additive change is visible even though the schema version holds. ## Known gaps Current as of 14 August 2026. These are the places where the format is weaker than the promise, listed so you can plan around them. **No content hash in the file.** Saive computes a SHA-256 over the canonical URL and normalised body, but it lives in the database, not the frontmatter. That means an outside tool can't currently tell whether a local file has diverged from what was captured. This belongs in `saive_extraction` and is the next change to this schema. **Images are referenced at their original URLs.** There is no attachment folder and no local copy. Nothing points at Saive's servers, but it does mean an exported library degrades as the web does, and isn't self-contained offline. Bundled attachments with relative paths, written at export, are planned. **There is no import path at all yet.** Not from other tools, and not from a Saive export. Saves currently enter Saive one at a time, through the extension or the web app. Bulk import is planned, and when it lands it will read a Saive export by matching on `uuid` first (same uuid means same save, updated in place, never duplicated), falling back to canonical URL for files that came from somewhere else. **`source` currently holds the domain**, not the human-readable site name the field name implies. `saive_source_metadata.domain` and `site_name` are defined but not yet populated. Read `source` as the domain until this note goes away. ## Changelog | Revision | Date | Change | |---|---|---| | 1.0 | 2026-08-14 | First public release. Describes `saive_schema: 1` as implemented. |