How Git Works Under the Hood

Most people learn Git as a pile of commands to memorize. Add this, commit that, and when something goes wrong, paste a scary incantation from a forum and hope. It feels confusing because the commands are taught before the model, and the model is actually simple. Git is a small content-addressed database[1] with a few object types and a couple of pointers on top. Learn that, and rebase, reset, and cherry-pick stop being magic.

Everything is stored by the hash of its content

The core idea is content addressing. When Git stores something, it names that thing by a hash of its own bytes. Identical content always produces the same hash, so the same file stored twice is only kept once. You can see this directly.

echo "hello" | git hash-object --stdin
# 3b18e512dba79e4c8300dd08aeb37f8e728b8dad

That 40 character string is a SHA-1 hash[2], and it is the object's address inside the repository. Git historically used SHA-1 and is moving toward SHA-256, but the principle is the same either way. Every object lives under .git/objects, filed by its hash. Change one byte of the content and you get a completely different hash, which is the property the entire system is built on.

Four object types

Git has only four kinds of objects, and they combine to represent your whole history.

A blob holds the raw contents of a file. A blob has no name and no path. It is just bytes and their hash. A tree represents a directory. It lists names, file modes, and the hash of the blob or sub-tree each name points to. The filenames live in the tree, not in the blob, which is why two files with identical content share one blob.

A commit is a snapshot. It points to one tree (the full state of your project at that moment), to its parent commit or commits, and it records the author, the committer, the timestamp, and the message. A tag object is an annotated, named pointer to another object, usually a commit, used for releases.

A commit is a snapshot, not a diff

This is the point that trips people up. A commit does not store "the lines you changed." It stores a pointer to a tree that describes the entire project as it looked then. Git shows you a diff by comparing two trees on the fly, but what it saves is a snapshot.

That sounds wasteful until you remember content addressing. If a file did not change between two commits, its blob hash is identical, so both trees point at the same existing blob. Unchanged files cost nothing to store again. You get the simplicity of snapshots with the storage efficiency of deduplication.

History is a graph

Because each commit records its parent, commits form a chain, and because a merge commit has two parents, that chain is really a directed acyclic graph. Directed because parent links point backward in time, and acyclic because you can never be your own ancestor.

A ─── B ─── C ─── E   (main)
             \     /
              D ──      (feature, merged at E)

When you read your project's history, you are reading this graph. git log --graph draws it for you. Operations that feel advanced are usually just walking or rewriting this graph.

Branches and HEAD are just pointers

Here is the part that makes branching feel cheap, because it is. A branch is not a copy of anything. It is a small file containing a single commit hash[3]. The main branch lives at .git/refs/heads/main and holds forty characters.

cat .git/refs/heads/main
# 7c3a9f1b2e...  (the commit main currently points to)

HEAD is one more pointer, usually pointing at the current branch rather than directly at a commit. When you commit, Git writes the new commit object and then updates your branch pointer to it. When you create a branch, Git writes one tiny file. When you switch branches, it moves HEAD and updates your working files to match. Nothing is duplicated, which is why creating and deleting branches is instant.

The staging area in the middle

Git has three places a file can live: your working directory, the repository, and a layer between them called the index, or staging area[4]. git add writes the file's blob into the object store and records it in the index. git commit takes the current index, builds a tree from it, and wraps that tree in a commit object.

This is why you stage changes before committing. The index lets you assemble exactly the snapshot you want, including only some of your edits, before you freeze it into history.

Why the model makes everything click

Once you hold these facts at the same time, snapshots addressed by content, history as a graph, branches as movable pointers, the confusing commands become obvious.

A merge finds a common ancestor in the graph and combines two lines of work into a commit with two parents. A rebase replays your commits one by one onto a different base, producing new commits with new hashes. A reset moves a branch pointer to a different commit. A cherry-pick takes the change from one commit and applies it as a new commit somewhere else. None of these are special cases. They are all moving pointers and creating objects in the same little database.

There is a security property in this too. Because each commit's hash includes its tree and its parent's hash, and that parent includes its own parent, the whole history is a chain of hashes. You cannot quietly alter an old commit without changing its hash, which changes every commit after it. Tampering is visible.

Git is not complicated. It is a content-addressed object store with four object types and a few pointers, and almost everything you do is composing those pieces. The commands make sense the moment you stop thinking in commands and start thinking in objects.

Sources (4)
  1. Pro Git: Git Internals - Git Objects
  2. Wikipedia: Git
  3. Pro Git: Git Internals - Git References
  4. Pro Git: Git Basics - Recording Changes to the Repository