Skip to content

ARCHITECTURE · 14

The VFS model.

TFS is tebako's userland virtual filesystem: a per-process mount table over bulk-storage images, with the kernel-VFS shape and none of the kernel. The design is one engine, one small C ABI, and pluggable backends, and the Rust tfs crate is the shipping implementation.

Note

The core, multi-mount, the tar adapter, the COW composite, and the ENC transform are shipped, and the C++ libtfs survives only as the parity oracle.

One engine, one seam

The product line is three things: libtfs, the engine, spoken through the tebako_fs_* C ABI; the tfs CLI, the human surface (tfs : libtfs :: sqlite3 : libsqlite3); and tebako-pkg, which does tpkg trailer surgery, a tebako concept rather than a TFS one. The ABI is additive-only (tebako_fs_abi_version() = 1), and exactly the tebako_* symbols are exported, nm-verified, with nothing else leaking.

Calls reach the router in three ways. The packaged interpreter’s patched IO calls the ABI in-process: open, pread, stat, opendir, and friends. Native executables get the VFS injected through the preload interposition shim (tfs exec), which is itself a TFS consumer. A dlopen of a VFS-resident library materializes the file to the exec cache first (tebako_fs_dlmap2file), so the OS loader maps a real file. Everything else, the working directory, writes outside held trees, mkdir, and unlink, goes to the host, policy-gated.

The mount table

Note

The mount table is shipped.

One process can attach many images at once: the env image at the runtime root, the app payload, and data payloads. Each mount gets a monotonic handle (never reused) and its own backend instance; there is no global mount state, which is what makes N concurrent mounts safe. Path resolution is longest-prefix dispatch on component boundaries: nested mounts shadow outer ones, and a duplicate mount point is refused with EEXIST. Every fd and dir handle records its owning mount, so unmounting one image force-closes only its own handles (later use fails EBADF) and leaves the rest fully usable.

open("/__app__/config/boot.rb", …)
        │
        ▼
┌───────────────────────────────────────────────────────┐
│  path dispatch — longest mount-point prefix wins      │
│                                                       │
│   mount 0   /__runtime__        → limnifs backend     │
│   mount 1   /__layers__/fonts   → dwarfs backend      │
│   mount 2   /__app__            → limnifs backend     │
│   (nested mount points allowed; duplicates → EEXIST)  │
└───────────────────────────────────────────────────────┘
        │ path inside a mount              │ path outside every mount
        ▼                                  ▼
  tebako_fs_open()                    host open()
  read-only image, owning mount       cwd, writes, mkdir — policy-gated
The TFS router: calls enter through the tebako_fs_* C ABI; the Rust TFS router dispatches by longest mount-point prefix onto read-only backends — limnifs, dwarfs-t, squashfs, tar, zip — with Rust COW and ENC transforms stacked above; paths outside every mount pass through to the host filesystem.

Figure 1 — The TFS router: longest-prefix dispatch onto read-only backends, with the COW and ENC transforms stacked above; everything else passes through to the host filesystem.

Symlinks play by the router rules too: relative targets resolve inside their own mount and never escape it, and absolute targets resolve through the mount table, so a link in one payload can land in another; cross-mount links work, and that is the point of composition. A target outside every mount is host-passthrough under the same jail policy as any host path, so a symlink can never widen a jail.

Coverage is not presence

A mount covers everything under its mount point, but its image holds only the entries it actually contains. A covered path the image does not hold does not surface the backend’s ENOENT; open, stat, and opendir fall through to the host answer, gated by the jail policy. With the app payload mounted at /, the v2 dispatch shape, this is exactly what keeps the host filesystem reachable: held content serves from the image (read-only), and everything else behaves as if no mount claimed it.

Note

The write gate, stated plainly. Writes into a held tree, a path the image holds or one with an existing in-image ancestor, are refused with EROFS. Zip images have no explicit directory entries, so implied parents count, and dwarfs always carries them. Covered-but-not-held paths pass to the host, policy-gated. There is no write surface behind a read-only mount and no way for a payload to corrupt its own image or another package’s mount.

Backend capabilities

Note

The backends are shipped, and limnifs is the default since v0.2.0.

Backends are pluggable drivers behind one trait, and new formats are additive. The capability model is honest per format, with no uniform read/write pretense. Runtime mounts are read-only by default (TEBAKO_MOUNT_RO); TEBAKO_MOUNT_COW stacks the composite transform, and TEBAKO_MOUNT_RW is ENOTSUP, because backends never learn to write.

FORMAT RUNTIME MOUNTS IN-PLACE EDITS CREATION NOTES

limnifs

read-only

—

in-process writer (limnifs-write)

It is the default since v0.2.0: pure Rust (#![forbid(unsafe_code)]), content-addressed (BLAKE3), LMFS magic, and it compiles everywhere Rust compiles.

dwarfs-t

read-only

—

in-process Writer (dwarfs-t-rs)

It is the only C++ in the stack, with FlatBuffers metadata that upstream dwarfs cannot read.

squashfs

read-only

—

mksquashfs-class tooling

It is POSIX-only, Windows builds ship without it, and a mount attempt fails ENOTSUP by name.

zip

read-only

add / delete (format layer)

—

It is pure Rust, it has no explicit directory entries, and implied parents count for the held check.

tar family

read-only

append-only

repack

It is pure Rust, gz / zst envelopes are detected by magic, and the tar heuristic always probes last.

Format detection keys on bytes, not on the trailer: strong magic first (zip PK\x03\x04, dwarfs, squashfs hsqs, limnifs LMFS), with the weak tar heuristic always last. The trailer’s format_id is a hint and answers exactly one question, how to read these bytes: 0 auto, 1 dwarfs, 2 squashfs, 3 zip, 5 limnifs. Id 4 stays the legacy runtime-role wart, never mounted and never reused. Runtime-role and entrypoint semantics live in the manifest, never in the format axis. A mount whose detected format has no compiled-in backend fails with the named ENOTSUP, never a silent re-route and never a partial read.

Transforms: COW and ENC

Note

The transforms machinery is shipped.

The declarative manifest surface is planned.

Transforms are not formats: they stack above a backend, and they exist only in the Rust TFS. The stack order is pinned: COW → ENC → format backend. COW is always outermost, and ENC decrypts exactly one image’s backend and sits underneath.

COW: the composite backend

CowBackend stacks a writable overlay over a read-only base: reads fall through to the base unless shadowed; writes, deletes, and attribute changes land in the overlay; and modifying a base file copies it up first, so the base image stays byte-identical. The overlay is a plain host directory, disposable by deleting it, and self-contained: a .tfs-whiteouts journal inside it records deleted paths (strict v1 text, rewritten atomically, and any malformed line fails the mount with EINVAL). Whiteouts mask base entries only, and an overlay entry of the same name always wins, which is overlayfs semantics. Layers stack, detach, ship, and re-attach.

ENC: the decrypting view

EncBackend wraps one image’s backend as a decrypting read view: directory structure, names, and symlink targets stay plaintext, and only regular-file content is ciphertext. A mount requires an opening grant resolved against the in-image envelope manifest, and a read of a still-sealed path answers ENOKEY (126), the named EKEY class, never garbage. Copy-up reads through it, so a writable sealed-base mount is COW(ENC(base)), and the plaintext lands in the operator-bound host store, outside the image’s confidentiality envelope. This point is stated plainly: protecting scratch at rest is the host’s layer (disk encryption), never a tebako declaration.

The declarative surface gates where writes may go: a slice declares write areas, the operator binds an overlay store, and a write outside every declared area stays EROFS, journaled as event=vfs-deny. Under record mode every image mount stacks an ephemeral scratch COW instead, and writes are journaled event=vfs-write, so tfs needs --from-journal can draft the declarations. Binding failures (an unbound retained store, a non-opening key, or a malformed TEBAKO_OVERLAYS / TEBAKO_DECRYPT) fail before exec with exit 68 (EX_TEBAKO_OVERLAY). Nothing is transformed that was not declared.

Mount sources: every file is a potential filesystem

Note

Mount-from-vfs is specified but not yet in the C ABI, and it is planned.

TFS is self-similar: images are files, and any file is mountable. Three mount sources are shipped: a host file, a file region (offset + length, which is how a package slot mounts without extraction), and memory. The fourth, the VFS-file-region, mounting an image addressed by a path inside an existing mount, with the new backend reading its bytes through the owning mount’s pread, is specified as the encapsulation primitive ("an FS includes another FS at a directory path") and is not yet surfaced in the C ABI. Nested mount points work today, and nested mount sources are the planned step.

The access matrix

Six ways to get at image contents exist, ordered by transparency. Each one is honest about what it costs and where it runs.

MECHANISM WHAT IT IS STATUS

link

in-process via libtfs — the runtime driver and the preload shim link the tfs crate directly

shipped; it is the fastest and deepest path, and it is the packaged-app model

fuse

tfs mount --fuse — a real system-wide mount wherever FUSE exists

planned; it is the one mechanism with a host dependency

serve

tfs mount --serve=nfs|webdav — a userland server the OS’s own client mounts

planned, with no FUSE anywhere

shell

tfs ls / tree / cat / stat / find / info — interactive browsing and one-shots

shipped; it is the human surface

exec

tfs exec image — cmd — the VFS injected through the preload interposition shim

shipped; it is the mainline native-exec mechanism, with macOS and linux-gnu first-class and Windows later

extract

tfs extract — materialize the image to disk

shipped; it is the honest fallback

Note

The trade is stated plainly: without the kernel there is no system-wide transparent mount for unmodified arbitrary processes. TFS trades that away deliberately, and it gets zero privileges, every platform, and in-process speed in return.

The debug contract

Note

The debug contract is shipped in crates/tebako-log.

One logging facility, tebako-log, serves the whole system, with every component wired through it; there is no ad-hoc eprintln! anywhere in the stack. Three environment variables control it:

VARIABLE MEANING

TEBAKO_DEBUG

It is the level: off (default) / error / warn / debug / trace, with per-component overrides such as debug,preload=trace,tfs=warn. The legacy boolean TEBAKO_DEBUG_TFS maps to debug.

TEBAKO_DEBUG_FILE

It is the sink: stderr by default, because stdout belongs to the payload and never to the log. A path opens append-mode with %p → pid expansion (exec’d children do not clobber one file); parents are created; and an unopenable path falls back to stderr with one warn line.

TEBAKO_DEBUG_COMPONENTS

It is a comma filter of component names (default: all): preload, tfs, driver, shim, bootstrap, cli, pkg, resolve.

tebako[84213] debug preload: route path=/x held=true action=erofs
   │        │      │         │      └─ one parseable line per event: <event> <k=v>…
   │        │      │         └─ component (a crate name)
   │        │      └─ level
   │        └─ pid
   └─ fixed prefix

The discipline is zero cost when off: the config is read once, and the gate runs before any formatting. Paths, decisions, and digests are loggable, while payload contents and key material never are.