File Systems

File Systems

Definition: An OS abstraction that organizes, stores, retrieves, and manages files on physical storage devices, translating human-facing paths and files into raw block-level reads and writes.

How It Works

  • Inodes: metadata structures storing file size, permissions, owner, timestamps, and pointers to the actual data blocks on disk. The filename itself is not stored in the inode, only in a directory entry that maps a name to an inode number.
  • Directory tables: map human-readable filenames to inode numbers. This is why hard links, multiple names pointing to the same inode, are possible, and why renaming a file doesn’t change its inode or its actual data location.
  • Journaling: logs pending metadata changes to a write-ahead log before committing them to their final location, so a crash mid-write can be replayed or rolled back on reboot instead of leaving the filesystem corrupted and unmountable.
  • Allocation strategies determine how a file’s data blocks are laid out on disk: contiguous allocation (fast sequential read, suffers external fragmentation), linked allocation (no fragmentation, poor random access since blocks must be followed as a chain), and indexed allocation via inode block pointers (used by ext-family filesystems, balances both).
  • Copy-on-Write filesystems (ZFS, Btrfs, APFS) never overwrite data in place. A modified block is written to a new location and metadata pointers are atomically updated, which is what enables instantaneous snapshots and eliminates a whole class of “torn write” corruption.

Under the Hood

An ext4 inode holds direct pointers to the first 12 data blocks of a file, plus indirect, doubly indirect, and triply indirect pointers for larger files, though modern ext4 mostly replaces this scheme with extents, a (start block, length) pair that describes a contiguous run of blocks in one entry instead of listing every block individually, which is both faster to traverse and cheaper to store for large, mostly-contiguous files.

The kernel’s Virtual File System (VFS) layer is what makes open(), read(), and write() work identically regardless of the underlying filesystem. VFS defines a common set of object types, superblock, inode, dentry (directory entry), and file, each with an operations table of function pointers. When a syscall reaches VFS, it dispatches through these function pointers into the specific filesystem driver (ext4, NTFS, a FUSE userspace filesystem, or an NFS client), which implements the actual block-level logic. This is also how the same read()/write() interface transparently works over a network filesystem.

Journaling in ext4 (via JBD2, the underlying journaling block device layer) can run in three modes: journal (both data and metadata are journaled, safest, slowest), ordered (only metadata is journaled, but data blocks are guaranteed to be written before their referencing metadata commits, the default), and writeback (only metadata is journaled with no ordering guarantee on data, fastest, riskiest). A crash during ordered mode can leave you with stale data in a block but never a metadata pointer to garbage.

Directory lookups themselves aren’t a linear scan of entries for any serious filesystem. ext4 indexes large directories with an HTree, a hashed B-tree keyed on a hash of each filename, so finding one file among millions in a single directory stays roughly logarithmic instead of degrading to a full directory scan. NTFS uses a similar B-tree indexing approach for its directory structures. This is why filesystems that predate these indexing schemes (old FAT variants) get dramatically slower as a single directory’s file count grows into the tens of thousands, while ext4/NTFS stay flat.

Debugging Workflow

“Disk full” and “why is this so slow” are the two most common filesystem-level production issues, and they’re diagnosed differently:

$ df -h          # space usage per mounted filesystem
$ df -i          # inode usage per mounted filesystem, catches inode exhaustion df -h misses
$ du -sh /var/log/*        # find which directory is actually consuming the space
$ lsof +L1                 # list open files with zero remaining links (deleted-but-still-open)

lsof +L1 specifically catches the “df says full but I deleted the logs” case: a deleted file held open by a running process (commonly a log file a service never rotated) still occupies disk space until that process closes it or exits, and lsof is the tool that reveals which process is the culprit.

$ iostat -x 1              # per-device I/O utilization, queue depth, await time
$ strace -T -e trace=read,write ./app     # time spent in individual read/write syscalls

A high %util or await in iostat output points at a genuine storage bottleneck rather than an application-level slowdown, useful for distinguishing “the disk is the problem” from “the code is doing something wasteful” before reaching for a filesystem-level fix.

Why It Matters

  • Guarantees data integrity, fast retrieval, access security via permission bits and ACLs, and crash safety on physical SSD/HDD media.
  • The filesystem layer is what every System Call like open(), read(), and write() ultimately routes through via VFS, which is why identical application code works uniformly across ext4, NTFS, and network filesystems.
  • SSD-aware filesystems must account for wear leveling and the fact that SSDs can’t overwrite in place at the byte level, they erase in large blocks first. Ignoring this leads to write amplification and premature drive wear, which is why TRIM/discard support matters for SSD longevity.

Common Pitfalls

  • Running out of inodes (allocating millions of tiny files) blocks new file creation even when disk space in bytes remains, since a fixed inode count is set at filesystem creation time. df shows space free while df -i reveals the real inode exhaustion.
  • Assuming write() returning success means data is durably on disk. Without an explicit fsync()/fdatasync() call, data can sit in the OS page cache and be lost on a power failure or kernel panic.
  • Deleting a large open file doesn’t actually free its disk space until every process holding it open closes its file descriptor, a classic cause of “df shows disk full but I deleted the log files” incidents.
  • Ignoring filesystem-specific limits (max path length, max file size, case sensitivity differences between ext4 and NTFS/APFS) causes bugs that only appear when code is ported across operating systems.
  • Assuming rename() is atomic across filesystems or volumes. It’s atomic within a single filesystem, but a rename across mount points/devices silently becomes a copy-then-delete, which is not atomic and can leave a partial file if interrupted.

History

  • Early Unix used simple, non-journaled filesystems (like the original FFS/UFS), where an unclean shutdown required a slow full-disk consistency check (fsck) on every boot to find and repair corruption.
  • Journaling filesystems (IBM’s JFS, SGI’s XFS, Linux’s ext3) appeared in the 1990s specifically to make crash recovery a fast, bounded-time journal replay instead of a full-disk scan, a huge win as disk sizes grew.
  • Copy-on-write filesystems (ZFS, 2005; Btrfs, first merged into Linux in 2009; APFS, 2017) represent the next major shift: eliminating in-place overwrites entirely removes torn-write corruption as a category and makes instant snapshots a natural side effect of the design rather than a bolted-on feature.

FAQ

Why does deleting a file not always free space immediately? Because the space is only reclaimed once the link count to the inode drops to zero and no process still has it open. A file with multiple hard links, or one held open by a running process, keeps its data alive until every reference is gone.

What’s the difference between formatting and partitioning? Partitioning divides a physical disk into logical sections; formatting writes an actual filesystem’s data structures (superblock, inode table, root directory) onto a partition so the OS can store files on it. A disk can be partitioned but unformatted, meaning no filesystem exists yet to write files into.

Does a filesystem need journaling to be reliable? Not strictly, copy-on-write filesystems achieve crash consistency without a traditional journal, by never overwriting live data in place. But a filesystem with neither journaling nor copy-on-write semantics (older FAT/UFS designs) risks structural corruption from an unclean shutdown.

Comparison

ext4NTFSAPFSZFS
Metadata structureJournaled inode tableMaster File Table (MFT)B-tree, copy-on-writeMerkle-tree of block pointers, copy-on-write
SnapshotsNo native supportVolume Shadow Copy (external)Native, instantNative, instant
Checksums on dataNo (metadata only via journal)NoMetadata onlyData and metadata
Typical platformLinuxWindowsmacOSSolaris, FreeBSD, Linux (via OpenZFS)
Max practical file size16TB256TB8EB16EB
Directory indexingHTree (hashed B-tree)B-treeB-treeMerkle-tree of block pointers

Filesystem Layout Quick Reference

ComponentStoresAnalogy
SuperblockFilesystem-wide metadata: size, block size, free block countThe filesystem’s “table of contents”
Inode tablePer-file metadata and data block pointersOne record per file/directory
Directory entriesFilename to inode number mappingsA phone book from names to inode numbers
Data blocksActual file contentThe pages of the book
JournalPending metadata changes not yet committedA scratch pad replayed after a crash

Example

Writing a file with the shell and tracing what the filesystem driver does:

$ echo "data" > file.txt
1. VFS routes open()/write() syscalls to the ext4 driver
2. ext4 allocates a data block, writes "data" into it
3. ext4 journal logs the metadata change (new inode, updated directory entry)
4. On commit, the journal entry is marked complete

If power is lost between steps 2 and 4, ext4’s journal replay on next mount can recover a consistent state instead of leaving a half-written, corrupted filesystem structure. Tools like stat file.txt show the inode number and metadata directly, and debugfs -R 'stat <inode>' /dev/sda1 can inspect raw ext4 inode structures for debugging.

Dig deeper