The Saive markdown schema
Every save in Saive is a markdown file with YAML frontmatter. Not a row in a database that can be exported to markdown, but a markdown file, written first, that a database happens to index.
This page is the specification for those files. It exists so that you can write your own tooling against your library, verify that an export is complete, or walk away entirely and still have something useful. It describes what the code actually emits today, not what's planned. Where the two differ, there's a known gaps section at the bottom that says so plainly.
The contract
- The file is authoritative. Every piece of per-save state Saive relies on round-trips through frontmatter and body. The database is an index, and rebuilding that index from the vault files is a real operation in the codebase, not a design aspiration. What a rebuild does not restore is listed in what doesn't travel.
- Identifiers are stable. A save's
uuidis assigned at creation and never changes, not on retitle, retag, move, or folder rename. It is also the filename. - Standard keys where standards exist. Obsidian's reserved keys (
tags,aliases), the Pandoc / Dublin Core overlap (title,author,description,source), and the common bookmarking keys (url,date_saved,status) are used as-is. - Extensions are namespaced. Anything Saive-specific is prefixed
saive_. Every one of those keys is safe to ignore, and safe to delete. - Nothing in a file points at Saive. No API endpoints, no callback URLs, no identifiers that only resolve against our servers. The two current exceptions are named in what doesn't travel.
File layout
A save is a single .md file named for its UUID:
Recipes/3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d.md
7e8f9a1b-2c3d-4e5f-6789-0abcdef12345.md
_concepts/edge-computing-c9d8.md
throughlines/0042-2026-07-04.md
Folders are real directories. A save with no folder sits at the root. Concept
files live under _concepts/, throughlines under throughlines/; both names are
reserved and can't be used as folder names.
Naming files by UUID rather than by title is deliberate: it means a bare
[[3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d]] wikilink resolves in Obsidian no
matter which folder the file has been moved to, and a retitle is never a rename.
Human-readable titles live in title and aliases, which is where Obsidian's
quick switcher looks anyway.
A complete file
---
tags:
- edge-compute
- infrastructure
aliases:
- The Quiet Revolution in Edge
title: The Quiet Revolution in Edge Compute
author: Jane Doe
description: >-
Edge computing has shifted from CDN supplement to a foundational layer of
cloud infrastructure.
source: example.com
url: https://example.com/2026/05/edge-compute
date_saved: "2026-05-17T14:32:00.000Z"
date_published: "2026-05-15T08:00:00.000Z"
status: unread
reading_time_min: 8
uuid: 3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d
created: "2026-05-17T14:32:00.000Z"
modified: "2026-05-17T14:33:12.000Z"
saive_schema: 1
saive_kind: article
saive_visibility: private
saive_summary: >-
The article traces three inflection points in edge computing and argues that
the edge is no longer a specialised tier but the centre of gravity for new
infrastructure.
saive_concepts:
- "[[Edge Computing]]"
saive_starred: true
saive_source_metadata:
og_image_url: https://example.com/img/edge.jpg
favicon_url: https://example.com/favicon.ico
saive_extraction:
tier: full
source: extension
paywall_detected: false
body_present: true
captured_at: "2026-05-17T14:32:00.000Z"
saive_processing:
pipeline_version: 1
ai_processed: true
skipped_reason: null
summary_model: claude-sonnet-4-6
tag_model: claude-haiku-4-5
processed_at: "2026-05-17T14:33:12.000Z"
---
# The Quiet Revolution in Edge Compute
For most of the last decade, "edge computing" meant a CDN with a few
lightweight execution primitives bolted on. ==That framing is now obsolete==[^hl-1]
## Notes <!-- saive:user-note -->
Worth revisiting when we plan next year's infra budget.
[^hl-1]: Compare against the 2024 latency numbers.
Field reference
Fields marked optional are omitted entirely when absent rather than written
as null or an empty array.
Obsidian-reserved
| Field | Required | Type | Notes |
|---|---|---|---|
tags | yes | string[] | Lowercase, hyphenated. May be empty. |
aliases | no | string[] | Alternative titles. When Saive rewrites an unusable captured title, the original is appended here. |
Standard interop
| Field | Required | Type | Notes |
|---|---|---|---|
title | yes | string | |
author | no | string | From the source page or AI extraction. |
description | no | string | The publisher's own one-or-two-line dek. Distinct from saive_summary. |
source | yes | string | Currently the domain. See known gaps. |
Bookmark fields
| Field | Required | Type | Notes |
|---|---|---|---|
url | yes | URL | The canonical URL of the saved page. |
date_saved | yes | ISO-8601 | When you saved it. |
date_published | no | string | The publisher's timestamp, stored verbatim as a string, not reformatted. |
status | yes | unread · read · archived | |
reading_time_min | no | integer | Estimate. |
Identity
| Field | Required | Type | Notes |
|---|---|---|---|
uuid | yes | UUIDv4 | Assigned at creation, never changes. Matches the filename. |
created | yes | ISO-8601 | Never changes. |
modified | yes | ISO-8601 | Updated on every write. |
Saive extensions
| Field | Required | Type | Notes |
|---|---|---|---|
saive_schema | yes | integer | Currently 1. |
saive_kind | yes | article · concept · throughline | Which of the three file shapes this is. |
saive_visibility | yes | private · public | Whether the save appears on a public folder feed. Absent is read as private. |
saive_shared_at | no | ISO-8601 datetime | When the save last turned public. Present only while saive_visibility is public. |
saive_summary | no | string | The AI summary: what the piece argues. |
saive_summary_descriptive | no | string | A neutral description of what the piece is. Used on shared surfaces where the substantive summary stays private to whoever saved it. |
saive_concepts | no | string[] | Obsidian wikilinks, e.g. "[[Edge Computing]]". These are what produce real edges in the graph view. |
saive_starred | no | true | Present only when starred. |
saive_user_edited_summary | no | true | Present only when you've edited the summary yourself, so a rebuild doesn't overwrite your words with a regenerated one. |
saive_source_metadata | no | object | og_image_url, og_image_caption, favicon_url. Emitted only when at least one is known. |
saive_extraction | yes | object | How the content was obtained. Below. |
saive_processing | yes | object | Whether AI ran, and what ran. Below. |
saive_archetype | no | object | What the content is. type holds the slug (recipe, video, …), with that archetype's data fields flat beneath it. Keys Saive doesn't recognise are preserved on read rather than dropped. |
saive_source | no | object | id + parser_version for the origin-specific parser that handled the save. |
saive_import | no | object | Present only on bulk-imported saves. format + session_id + imported_at. Below. |
saive_import
Present only on saves that arrived through bulk import, and absent on everything you capture with the extension, the web app, or the MCP server.
| Field | Required | Type | Notes |
|---|---|---|---|
format | yes | netscape · instapaper · pocket · raindrop · readwise | The export the save came from. |
session_id | yes | UUID | The import run. It means nothing outside Saive; safe to delete, safe to ignore. |
imported_at | yes | ISO-8601 | When Saive ingested the file. |
The timestamp rule is the part worth knowing. On an imported save,
created and date_saved carry the source's saved-at date — a link you
saved in Pocket in 2014 still reads as 2014, because the history is most of
what a library is worth. imported_at is the only field that records when the
file actually arrived. It's also how Saive keeps a five-thousand-item import
from burying everything you saved last week.
This block is orthogonal to saive_source below, which records which parser
extracted the content: an imported YouTube link carries both.
saive_extraction
Records how the body in this file was obtained, so a thin save is distinguishable from a real capture without guessing.
| Field | Required | Type | Notes |
|---|---|---|---|
tier | yes | full · preview · metadata_only · url_only | Quality of the capture. |
source | yes | extension · server | Browser extension capture, or server-side fetch. |
paywall_detected | yes | boolean | |
body_present | yes | boolean | Whether the article body is actually in this file. Must agree with the file; a mismatch is a corruption signal. |
captured_at | yes | ISO-8601 | |
format | no | pdf | Present when the body was converted from another document format. Its absence means ordinary HTML extraction. |
saive_processing
| Field | Required | Type | Notes |
|---|---|---|---|
pipeline_version | yes | integer | |
ai_processed | yes | boolean | Whether AI ran and produced metadata for this save. |
skipped_reason | yes | over_quota · absent · failed · null | null when ai_processed is true. |
summary_model | when processed | string | |
tag_model | when processed | string | |
processed_at | when processed | ISO-8601 | |
title_model | no | string | Set when AI replaced an unusable captured title. Independent of ai_processed, so a link-only save can get a title fix without any summary. |
title_source | no | pdf_heading · pdf_filename | Low-confidence provenance for a heuristically derived PDF title. |
embedding_ref | no | string | A pointer into Saive's search index. See what doesn't travel. |
archetype_confidence | no | number 0–1 | |
archetype_detected_by | no | schema_org · og · url_pattern · llm · user | user is never overwritten by automated detection. |
Body conventions
The body is the extracted article, converted to markdown with its structure intact: headings, lists, code blocks, blockquotes, tables.
H1. The first line is an H1 matching title.
Highlights are Obsidian-flavoured ==highlight== markers inline in the
prose. A highlight that carries a note gets a footnote reference immediately
after the closing ==, in the form [^hl-1], with its definition at the end of
the file. Publisher footnotes ([^1], [^note]) are left alone; only the
hl- prefix is ours. Highlights are not stored in frontmatter; the marked-up
body is the storage.
Your notes live under a heading with an explicit sentinel:
## Notes <!-- saive:user-note -->
The sentinel is the only rule. There's no heuristic. A ## Notes heading
without it is treated as article content, so an article that happens to have a
"Notes" section is never mistaken for your annotation. Everything after the
sentinel is yours; Saive's AI never rewrites it.
Concept files
Concepts are clusters Saive derives across your saves, one file each under
_concepts/, with saive_kind: concept. Each has two marked regions:
## Contributing saves
<!-- saive:auto-managed:start -->
- [[3a4b5c6d-7e8f-4a1b-9c2d-1e2f3a4b5c6d]]
<!-- saive:auto-managed:end -->
## Your notes
<!-- saive:user-block:start -->
Anything you write here survives regeneration.
<!-- saive:user-block:end -->
Only the text between the auto-managed sentinels is ever rewritten. Content outside it, including but not limited to the user block, is preserved verbatim. If the auto-managed sentinels are missing from a file, Saive refuses to guess where the list belongs rather than overwriting your edits.
The wikilinks are what make this useful outside Saive: drop the folder into Obsidian and the graph view shows articles connected to the concepts they contributed to, with no plugin and no import step.
Throughlines
Periodic AI syntheses of recent saves, one file per issue under
throughlines/, with saive_kind: throughline. Alongside the identity fields
they carry issue, published, coverage_start / coverage_end, themes,
generated_by, and a sources list mapping each paragraph of the narrative
back to the UUIDs it drew on. Always written saive_visibility: private.
Reading rules
If you're writing something that consumes these files:
Parse leniently. Saive's own reader does. Unknown keys, saive_-prefixed
or not, are preserved, not dropped. An unrecognised value in a constrained
field degrades to "unset" rather than throwing.
Identity is the only hard requirement. uuid, title, url, source,
created, status, and tags must be present and well-formed. Everything else
has a defined fallback.
Accept two shapes for every timestamp. Saive writes them quoted, so YAML
reads them back as strings. Obsidian's property editor rewrites frontmatter and
strips the quotes, after which the same values parse as YAML timestamps:
Date objects under js-yaml and PyYAML, not strings. This is expected, not
corruption. Handle both. Along the same lines: don't assume a file that has
passed through another editor is byte-identical to what Saive wrote. null
becomes blank, folded scalars reflow.
Wikilinks may lose their brackets. saive_concepts entries are written as
"[[Label]]", but a hand-edited file may carry a bare Label. Accept both.
What doesn't travel
Some things are honestly not portable, and it's better to name them than to let you find out:
saive_processing.embedding_ref is a pointer into Saive's vector index. It
means nothing outside Saive and recomputing it requires re-embedding the text.
It is safe to delete and safe to ignore.
saive_source.id names a parser: the code that handled a YouTube page or
an arXiv paper differently from a generic article. Archetype definitions are
declarative data and travel fine, but parsers are executable code and don't
serialise into a markdown file. Outside Saive that field is an opaque string.
Sharing state isn't in the files at all. Folder sharing levels, public URLs, feed tokens and invited-reader grants live only in the index. If the index is rebuilt from vault files, folders come back private and member-less. That fails closed by design: the worst case is a dead share link, never an accidental exposure. Empty folders disappear in the same way, since a folder in the vault is only ever a directory with saves in it.
Everything else in the file is either self-describing or standard.
Versioning
saive_schema is an integer, currently 1. Additive optional fields don't bump
it; the readers that matter, including any you write, should ignore fields
they don't know. It bumps only when the meaning of an existing field changes, and
that comes with a migration. A field's meaning is never silently repurposed.
The revision of this document is tracked separately, in the changelog below, so an additive change is visible even though the schema version holds.
Known gaps
Current as of 14 August 2026. These are the places where the format is weaker than the promise, listed so you can plan around them.
No content hash in the file. Saive computes a SHA-256 over the canonical URL
and normalised body, but it lives in the database, not the frontmatter. That
means an outside tool can't currently tell whether a local file has diverged
from what was captured. This belongs in saive_extraction and is the next
change to this schema.
Images are referenced at their original URLs. There is no attachment folder and no local copy. Nothing points at Saive's servers, but it does mean an exported library degrades as the web does, and isn't self-contained offline. Bundled attachments with relative paths, written at export, are planned.
There is no import path at all yet. Not from other tools, and not from a
Saive export. Saves currently enter Saive one at a time, through the extension
or the web app. Bulk import is planned, and when it lands it will read a Saive
export by matching on uuid first (same uuid means same save, updated in place,
never duplicated), falling back to canonical URL for files that came from
somewhere else.
source currently holds the domain, not the human-readable site name the
field name implies. saive_source_metadata.domain and site_name are defined
but not yet populated. Read source as the domain until this note goes away.
Changelog
| Revision | Date | Change |
|---|---|---|
| 1.0 | 2026-08-14 | First public release. Describes saive_schema: 1 as implemented. |