Computing and the Command Line

Why Version Control, and How Git Thinks

Version control records every saved state of a project so you can see what changed, go back, and work on several things at once. Git, written in 2005 for the Linux kernel, stores history as snapshots of the whole project rather than lists of changes, keeps the full history on your own machine, and names every snapshot by a checksum. The three states a file can be in, and the three areas they live in.

  • 5 min
  • 7 steps
  • 2 questions
  • Lesson 25 of 80

In this lesson

  1. What version control is for
  2. Snapshots, not differences
  3. Everything is local
  4. Integrity
  5. Three states, three areas
  6. Your turn
  7. So

What version control is for

You already do version control by hand: budget-final.ods, budget-final-2.ods, budget-REALLY-final.ods. A version control system does it properly. It records each saved state of a set of files, with who saved it, when, and why, so you can 1 2:

  • see exactly what changed between any two versions,
  • go back to any earlier state, of one file or the whole project,
  • work on several things at once, then combine them,
  • and share work with other people, or with your other computers, without overwriting each other.

It’s essential for software, and just as useful for scripts, configuration files, notes, and anything else made of text.

Git is the version control system nearly everyone uses. Linus Torvalds and the Linux kernel developers wrote it in 2005, after they lost free use of the commercial tool they’d been using 1.

Snapshots, not differences

Many older systems stored a file’s history as a list of changes. Git thinks differently: each time you save, a commit, it records a snapshot of what every tracked file looks like at that moment 1. If a file hasn’t changed, git doesn’t store it again; the new snapshot simply points to the copy it already has 1.

Each snapshot also remembers its parent, the snapshot before it, so history is a chain of snapshots, and later a branching graph of them 2.

Three commits in a row, each labeled with a short ID and message: adc6083 Start the garden plan; 4e1b2a9 Add carrots; 9c3d7f0 Plan the fence, each pointing back to its parent. Under each are the files in that snapshot: README.md and beds.txt version 1; README.md and beds.txt version 2; README.md, beds.txt version 2, and fence.txt. Boxes of the same color are the very same stored copy, so only changed files are stored again. A main label points at the newest commit and HEAD points at main. Below, a commit's full ID, adc608396be93380956acec8ccb105797b4a2e75, is a SHA-1 hash of its contents: change anything and the ID changes, and the first seven characters are usually enough to name it.
History as a chain of snapshots, each named by its own checksum. Credit: StudyCorner diagram · CC BY 4.0 · Source

Quick check

How does git store the history of a project?

Everything is local

A git repository holds the project’s entire history, and it lives right in your project folder, in a hidden subfolder named .git 1. Looking at old versions, comparing them, making commits, and creating branches all happen on your own disk, so they’re instant and work offline 1.

Services like GitHub host a copy of a repository so you can share it and back it up. But every copy, including yours, is a full repository; GitHub isn’t where git “lives” 1.

Quick check

Can you look at old versions and make commits with no internet connection?

Integrity

Everything git stores is checksummed and referred to by that checksum: a SHA-1 hash, 40 hexadecimal characters like 24b9da6552252987aa493b52f8696cd6d3b00373, computed from the contents 1. Change one byte of one file and the hash changes, so git notices corruption, and history can’t be quietly altered. You’ll see these IDs everywhere; the first seven characters are usually enough to name a commit.

Once something is committed, it’s very hard to lose, especially if the repository is also copied somewhere else 1. Uncommitted work is another story, which is a good reason to commit often.

Three states, three areas

The one thing Pro Git asks you to remember above all: a tracked file is always in one of three states 1:

  • Modified: you’ve changed it, but haven’t marked it to be saved.
  • Staged: you’ve marked this version of it to go into the next commit.
  • Committed: it’s safely stored in the repository.

Those states correspond to three areas 1:

  • the working directory, the files you see and edit;
  • the staging area (also called the index), a list of what will go into the next commit;
  • the repository, the .git folder, where commits live.

So the basic rhythm is: edit files, stage the changes you want with git add, and commit them with git commit. The staging area is what lets you commit just part of what you’ve changed, which you’ll use constantly. Lesson 3 does all of this for real.

Your turn

Think it through

  1. List three things in your own files that would benefit from version control.
  2. In the diagram, why does the third commit take almost no extra space, even though it shows three files?
  3. You edited beds.txt and ran git add beds.txt, then edited it again. What state is the file in now?
Answers
  1. For example: the scripts and config files from the shell course, your ~/.bashrc and other settings, a house-project notes folder, writing in plain text or Markdown.
  2. README.md and beds.txt are unchanged, so the snapshot points to copies git already has; only fence.txt is new.
  3. Both: the first edit is staged, and the second is a modification in the working directory that isn’t staged yet. git status shows the file in both lists.

So

Version control records every saved state of a project. Git stores each commit as a snapshot of the whole project, sharing unchanged files, chained to its parent and named by a SHA-1 checksum. The full history lives in your project’s .git folder, and every file is modified, staged, or committed: working directory, staging area, repository.

Lesson complete

Nice work.

1day streak
0/1today's goal
–correct

Up next · 4 min

Installing and Configuring Git

Next lesson
Sources for this lesson
  1. 1
    Scott Chacon, Ben Straub. Pro Git, 2nd edition. Apress; free online at git-scm.com. 2014. verifiedFree CC BY-NC-SA 3.0 book, maintained online. Ch. 1: version control; Git's 2005 origin when the Linux kernel lost free use of BitKeeper; snapshots, not differences (unchanged files stored once); nearly every operation local; integrity through 40-character SHA-1 checksums; the three states (modified, staged, committed) and three areas (working tree, staging area or index, .git directory); first-time setup with system/global/local config levels, user.name and user.email baked into commits, core.editor, git config --list --show-origin. Ch. 2: git init, status (and -s), add, diff and diff --staged, commit (-m, -a), .gitignore, log options, amending, undoing, remotes, tags, aliases. Ch. 3: branches as movable pointers, HEAD, merging and conflicts, remote branches, rebasing and its rule. Ch. 7: reset demystified, stashing, revision selection. Ch. 8: core.autocrlf true on Windows, input on Linux and macOS. Ch. 10: objects (blob, tree, commit) and references.
  2. 2
    Anish Athalye, Jon Gjengset, Jose Javier Gonzalez Ortiz. Version Control and Git (The Missing Semester of Your CS Education, 2026). MIT CSAIL. 2026. verifiedCC BY-NC-SA. Explains git from its data model up: a version control system tracks snapshots of a top-level directory with metadata; in git, files are blobs, directories are trees, and history is a directed acyclic graph of snapshots, each pointing to its parents; objects are content-addressed by hash and references give them human-readable names; the staging area chooses what goes into the next snapshot.