How DNS Actually Works
Before your browser can load a single byte of a website, it has to answer a question: what number do I connect to? People use names like example.com, but machines route traffic to IP addresses like 93.184.216.34. The system that translates between the two is DNS, the Domain Name System[1], and it handles this for every name on the internet, for billions of lookups a second, without any single computer being in charge. It is worth understanding because when it breaks, everything looks broken, and the fix is usually simple once you know the shape of the thing.
The problem it solves
Humans are good at names and bad at numbers. Machines are the opposite. You could keep a giant file mapping every name to every address, and in fact the early internet did exactly that with a shared text file[2]. That stopped working almost immediately, because the list grew too fast and everyone needed to agree on one copy. DNS replaced the single list with a hierarchy that spreads responsibility across millions of servers and leans heavily on caching.
A hierarchy you read right to left
A domain name is a path through a tree, and you read it from the right. Take www.example.com with its invisible trailing dot, www.example.com.. The dot at the end is the root[3]. To the left of it is the top-level domain[4], com. To the left of that is example, and www is a name within that.
Each level is responsible only for delegating the level below it. The root knows where to find the servers for com. The com servers know where to find the servers for example.com. The example.com servers know the actual address of www. No server has to know everything. Each one only has to know who to ask next.
Who actually answers
Three kinds of authoritative servers form that hierarchy. The root servers sit at the top[5] and point you toward the right top-level domain servers. The top-level domain servers, for com, org, io, and the rest, point you toward the servers responsible for a specific domain. The authoritative servers for a domain hold the real records, the actual addresses and configuration that someone set up for that name.
Doing the walking is a separate job, handled by a recursive resolver. This is usually run by your internet provider, or a public one like 1.1.1.1 or 8.8.8.8[6]. Your device asks the resolver a simple question and expects a final answer. The resolver does the legwork of asking the hierarchy.
A lookup, step by step
Say your resolver has nothing cached and needs to find www.example.com from scratch.
your device -> resolver: what is www.example.com?
resolver -> root: who handles .com?
root -> resolver: ask the .com servers (here they are)
resolver -> .com: who handles example.com?
.com -> resolver: ask example.com's servers (here they are)
resolver -> example.com: what is www?
example.com -> resolver: 93.184.216.34
resolver -> your device: 93.184.216.34
That full walk looks like a lot of back and forth, and it is, which is why it almost never happens. Most of the time the resolver already has the answer cached and replies in a millisecond.
The records themselves
A domain is configured with a set of records, each a different type for a different purpose.
An A record maps a name to an IPv4 address[7]. An AAAA record maps a name to an IPv6 address. A CNAME record is an alias, pointing one name at another name[8] instead of an address, which is how www.example.com often points at a provider's hostname. An MX record says which servers receive email for the domain[9]. A TXT record holds arbitrary text, used for things like verifying domain ownership and email authentication. An NS record names the authoritative servers for the domain[10], which is the delegation itself.
Caching and the TTL
Every record carries a time to live, a TTL, measured in seconds[11]. When a resolver gets an answer, it keeps it for that long before asking again. TTLs are why DNS scales. The root and top-level servers would collapse instantly if every lookup went all the way up the tree, but caching means the vast majority of questions are answered close to the user.
This is also the source of the thing people call propagation. When you change a record, the new value is live at the authoritative server right away, but resolvers around the world keep serving the old cached value until its TTL expires. There is no global push. Everyone is just waiting out their caches. The practical trick before a planned change is to lower the TTL a day ahead of time, so the window of stale answers afterward is short.
Why it rarely breaks, and when it does
DNS is robust because it is both cached and redundant. Every level has multiple servers in multiple places, so the failure of any one is invisible. When something does go wrong, the cause is usually human. A record points at the wrong place, a TTL was set to a day so a fix takes a day to take hold, or a domain's delegation was never updated after a move. Because so much depends on it, a small DNS mistake can take down services that are otherwise perfectly healthy, which is why "it was DNS" has become an engineering punchline.
The tired description of DNS is a phone book. A better one is a chain of delegation. No one holds the whole directory. Each server knows only who to ask next, and caching makes that chain feel instant. It is a design that has scaled from a few hundred computers to the entire internet without changing its fundamental shape, which is about the highest compliment you can pay a piece of infrastructure.
Sources (11)
- Wikipedia: Domain Name System
- Wikipedia: Hosts (file)
- Wikipedia: DNS root zone
- Wikipedia: Top-level domain
- Wikipedia: Root name server
- Wikipedia: Public recursive name server
- RFC 1035: Domain Names - Implementation and Specification
- RFC 1034: Domain Names - Concepts and Facilities
- RFC 5321: Simple Mail Transfer Protocol
- RFC 1034: Domain Names - Concepts and Facilities
- RFC 1035: Domain Names - Implementation and Specification