open("/__app__/config/boot.rb", …)
│
▼
┌───────────────────────────────────────────────────────┐
│ path dispatch — longest mount-point prefix wins │
│ │
│ mount 0 /__runtime__ → limnifs backend │
│ mount 1 /__layers__/fonts → dwarfs backend │
│ mount 2 /__app__ → limnifs backend │
│ (nested mount points allowed; duplicates → EEXIST) │
└───────────────────────────────────────────────────────┘
│ path inside a mount │ path outside every mount
▼ ▼
tebako_fs_open() host open()
read-only image, owning mount cwd, writes, mkdir — policy-gated
ARCHITECTURE · 14
The VFS model.
TFS is tebako's userland virtual filesystem: a per-process mount table over bulk-storage images, with the kernel-VFS shape and none of the kernel. The design is one engine, one small C ABI, and pluggable backends, and the Rust tfs crate is the shipping implementation.
|
Note
|
The core, multi-mount, the tar adapter, the COW composite, and the ENC transform are shipped, and the C++ libtfs survives only as the parity oracle. |
One engine, one seam
The product line is three things: libtfs, the engine, spoken through the
tebako_fs_* C ABI; the tfs CLI, the human surface (tfs : libtfs ::
sqlite3 : libsqlite3); and tebako-pkg, which does tpkg trailer surgery, a
tebako concept rather than a TFS one. The ABI is additive-only
(tebako_fs_abi_version() = 1), and exactly the tebako_* symbols are
exported, nm-verified, with nothing else leaking.
Calls reach the router in three ways. The packaged interpreter’s patched IO
calls the ABI in-process: open, pread, stat, opendir, and friends. Native
executables get the VFS injected through the preload interposition shim
(tfs exec), which is itself a TFS consumer. A dlopen of a VFS-resident
library materializes the file to the exec cache first
(tebako_fs_dlmap2file), so the OS loader maps a real file. Everything
else, the working directory, writes outside held trees, mkdir, and unlink,
goes to the host, policy-gated.
The mount table
|
Note
|
The mount table is shipped. |
One process can attach many images at once: the env image at the runtime
root, the app payload, and data payloads. Each mount gets a monotonic handle
(never reused) and its own backend instance; there is no global mount state,
which is what makes N concurrent mounts safe. Path resolution is
longest-prefix dispatch on component boundaries: nested mounts shadow
outer ones, and a duplicate mount point is refused with EEXIST. Every fd
and dir handle records its owning mount, so unmounting one image
force-closes only its own handles (later use fails EBADF) and leaves the
rest fully usable.
Figure 1 — The TFS router: longest-prefix dispatch onto read-only backends, with the COW and ENC transforms stacked above; everything else passes through to the host filesystem.
Symlinks play by the router rules too: relative targets resolve inside their own mount and never escape it, and absolute targets resolve through the mount table, so a link in one payload can land in another; cross-mount links work, and that is the point of composition. A target outside every mount is host-passthrough under the same jail policy as any host path, so a symlink can never widen a jail.
Coverage is not presence
A mount covers everything under its mount point, but its image holds
only the entries it actually contains. A covered path the image does not
hold does not surface the backend’s ENOENT; open, stat, and opendir fall
through to the host answer, gated by the jail policy. With the app payload
mounted at /, the v2 dispatch shape, this is exactly what keeps the host
filesystem reachable: held content serves from the image (read-only), and
everything else behaves as if no mount claimed it.
|
Note
|
The write gate, stated plainly. Writes into a held tree, a path the
image holds or one with an existing in-image ancestor, are refused with
|
Backend capabilities
|
Note
|
The backends are shipped, and limnifs is the default since v0.2.0. |
Backends are pluggable drivers behind one trait, and new formats are
additive. The capability model is honest per format, with no uniform
read/write pretense. Runtime mounts are read-only by default
(TEBAKO_MOUNT_RO); TEBAKO_MOUNT_COW stacks the composite transform, and
TEBAKO_MOUNT_RW is ENOTSUP, because backends never learn to write.
| FORMAT | RUNTIME MOUNTS | IN-PLACE EDITS | CREATION | NOTES |
|---|---|---|---|---|
|
read-only |
— |
in-process writer (limnifs-write) |
It is the default since v0.2.0: pure Rust ( |
|
read-only |
— |
in-process Writer (dwarfs-t-rs) |
It is the only C++ in the stack, with FlatBuffers metadata that upstream dwarfs cannot read. |
|
read-only |
— |
mksquashfs-class tooling |
It is POSIX-only, Windows builds ship without it, and a mount attempt fails ENOTSUP by name. |
|
read-only |
add / delete (format layer) |
— |
It is pure Rust, it has no explicit directory entries, and implied parents count for the held check. |
|
read-only |
append-only |
repack |
It is pure Rust, gz / zst envelopes are detected by magic, and the tar heuristic always probes last. |
Format detection keys on bytes, not on the trailer: strong magic first (zip
PK\x03\x04, dwarfs, squashfs hsqs, limnifs LMFS), with the weak tar
heuristic always last. The trailer’s format_id is a hint and answers
exactly one question, how to read these bytes: 0 auto, 1 dwarfs, 2 squashfs,
3 zip, 5 limnifs. Id 4 stays the legacy runtime-role wart, never mounted and
never reused. Runtime-role and entrypoint semantics live in the manifest,
never in the format axis. A mount whose detected format has no compiled-in
backend fails with the named ENOTSUP, never a silent re-route and never a
partial read.
Transforms: COW and ENC
|
Note
|
The transforms machinery is shipped. The declarative manifest surface is planned. |
Transforms are not formats: they stack above a backend, and they exist only
in the Rust TFS. The stack order is pinned: COW → ENC → format backend.
COW is always outermost, and ENC decrypts exactly one image’s backend and
sits underneath.
COW: the composite backend
CowBackend stacks a writable overlay over a read-only base: reads fall
through to the base unless shadowed; writes, deletes, and attribute changes
land in the overlay; and modifying a base file copies it up first, so the
base image stays byte-identical. The overlay is a plain host directory,
disposable by deleting it, and self-contained: a .tfs-whiteouts journal
inside it records deleted paths (strict v1 text, rewritten atomically, and
any malformed line fails the mount with EINVAL). Whiteouts mask base
entries only, and an overlay entry of the same name always wins, which is
overlayfs semantics. Layers stack, detach, ship, and re-attach.
ENC: the decrypting view
EncBackend wraps one image’s backend as a decrypting read view: directory
structure, names, and symlink targets stay plaintext, and only regular-file
content is ciphertext. A mount requires an opening grant resolved against
the in-image envelope manifest, and a read of a still-sealed path answers
ENOKEY (126), the named EKEY class, never garbage. Copy-up reads through
it, so a writable sealed-base mount is COW(ENC(base)), and the plaintext
lands in the operator-bound host store, outside the image’s confidentiality
envelope. This point is stated plainly: protecting scratch at rest is the
host’s layer (disk encryption), never a tebako declaration.
The declarative surface gates where writes may go: a slice declares write
areas, the operator binds an overlay store, and a write outside every
declared area stays EROFS, journaled as event=vfs-deny. Under record
mode every image mount stacks an ephemeral scratch COW instead, and writes
are journaled event=vfs-write, so tfs needs --from-journal can draft the
declarations. Binding failures (an unbound retained store, a non-opening
key, or a malformed TEBAKO_OVERLAYS / TEBAKO_DECRYPT) fail before exec
with exit 68 (EX_TEBAKO_OVERLAY). Nothing is transformed that was not
declared.
Mount sources: every file is a potential filesystem
|
Note
|
Mount-from-vfs is specified but not yet in the C ABI, and it is planned. |
TFS is self-similar: images are files, and any file is mountable. Three mount sources are shipped: a host file, a file region (offset + length, which is how a package slot mounts without extraction), and memory. The fourth, the VFS-file-region, mounting an image addressed by a path inside an existing mount, with the new backend reading its bytes through the owning mount’s pread, is specified as the encapsulation primitive ("an FS includes another FS at a directory path") and is not yet surfaced in the C ABI. Nested mount points work today, and nested mount sources are the planned step.
The access matrix
Six ways to get at image contents exist, ordered by transparency. Each one is honest about what it costs and where it runs.
| MECHANISM | WHAT IT IS | STATUS |
|---|---|---|
|
in-process via libtfs — the runtime driver and the preload shim link the tfs crate directly |
shipped; it is the fastest and deepest path, and it is the packaged-app model |
|
|
planned; it is the one mechanism with a host dependency |
|
|
planned, with no FUSE anywhere |
|
|
shipped; it is the human surface |
|
|
shipped; it is the mainline native-exec mechanism, with macOS and linux-gnu first-class and Windows later |
|
|
shipped; it is the honest fallback |
|
Note
|
The trade is stated plainly: without the kernel there is no system-wide transparent mount for unmodified arbitrary processes. TFS trades that away deliberately, and it gets zero privileges, every platform, and in-process speed in return. |
The debug contract
|
Note
|
The debug contract is shipped in crates/tebako-log. |
One logging facility,
tebako-log,
serves the whole system, with every component wired through it; there is no
ad-hoc eprintln! anywhere in the stack. Three environment variables
control it:
| VARIABLE | MEANING |
|---|---|
|
It is the level: off (default) / error / warn / debug / trace, with per-component overrides such as debug,preload=trace,tfs=warn. The legacy boolean TEBAKO_DEBUG_TFS maps to debug. |
|
It is the sink: stderr by default, because stdout belongs to the payload and never to the log. A path opens append-mode with %p → pid expansion (exec’d children do not clobber one file); parents are created; and an unopenable path falls back to stderr with one warn line. |
|
It is a comma filter of component names (default: all): preload, tfs, driver, shim, bootstrap, cli, pkg, resolve. |
tebako[84213] debug preload: route path=/x held=true action=erofs │ │ │ │ └─ one parseable line per event: <event> <k=v>… │ │ │ └─ component (a crate name) │ │ └─ level │ └─ pid └─ fixed prefix
The discipline is zero cost when off: the config is read once, and the gate runs before any formatting. Paths, decisions, and digests are loggable, while payload contents and key material never are.