Skip to main content

Command Palette

Search for a command to run...

How do web browsers work under the hood?

Published
•11 min read•View as Markdown
N

Neha Chopra | Finance Writer → Full-Stack Developer in Progress I write finance case studies and analytical articles, and I’m currently transitioning into web development. Passionate about building clean, functional web experiences while bringing an analytical mindset from finance into tech.

We all have used web browsers like chrome and edge for research, shopping, consuming content and for many other uses. But we rarely know what actually happens under the hood. A web browser is far more than just a window to the internet. It’s a sophisticated piece of software that acts as interpreter, translator and renderer, transforming code into the rich, interactive experiences we see on our screens.

What is a browser?

At its core, a web browser is s software application that retrieves, interprets and displays content from world wide web. It is someone that speaks multiple languages (js, html, css), understands various protocols (http, https, webSocket) and paints pixels on our screen to create everything from simple text pages to complex web applications.

Beyond “Opening websites”, browser manages security(guarding us against malicious sites), handles user data (cookies, passwords, bookmarks), executes code and provides a platform for developers to build sophisticated applications. Modern browsers are essentially operating systems within operating systems, capable of running games, video editors, and productivity suites that would have required dedicated software just a decade ago.

Main components of a browser

A web browser is built with several major components working in concert, each with specific responsibility! These are

Components of a web browser

Components of a web browser

1. User Interface (UI)

This is the part we see and interact with directly: the address bar where we type the urls, the back and forward buttons, the refresh button, bookmarks and tabs. The UI is everything except the actual web page content. When we click back buttons or open a new tab, we are interacting with user’s UI layer.

2. Browser Engine

The browser engine acts as a bridge between the user interface and the rendering engine. When we type a URL in the address bar and hit Enter, the browser engine takes that action and orchestrates what happens next. It manages the rendering engine, handles navigation, manages sessions, and processes high-level browser operations.

3. Rendering Engine

This is where the magic happens. The rendering engine (sometimes called the layout engine) is responsible for displaying the requested content. If we're requesting HTML, the rendering engine parses the HTML and CSS, constructs a visual representation, and displays it on our screen. Different browsers use different rendering engines: Chrome and Edge use Blink, Firefox uses Gecko, and Safari uses WebKit. The rendering engine is the artist that transforms code into pixels.

4. Networking Component

The networking layer handles all internet communication. It implements protocols like HTTP and HTTPS, manages requests and responses, handles caching, and deals with cookies. When we click a link, the networking component fetches the HTML document from the server, along with all the associated resources like images, stylesheets, and scripts. It handles everything from DNS lookups to establishing secure connections.

5. JavaScript Engine

Modern websites aren't just static documents—they're interactive applications, and that interactivity comes from JavaScript. The JavaScript engine (like V8 in Chrome or SpiderMonkey in Firefox) interprets and executes JavaScript code.

6. UI Backend

This component provides basic drawing and windowing capabilities using the operating system's user interface methods. It draws basic widgets like combo boxes, windows, and checkboxes. The UI backend ensures that the browser looks and feels appropriate for the operating system it's running on.

7. Data Persistence Layer

Browsers need to remember things: our bookmarks, our browsing history, cookies, cached files, and saved passwords. The data persistence layer manages all of this local storage. It implements storage mechanisms like localStorage, IndexedDB, and the browser cache, allowing websites to save data on our computer and retrieve it later.

Networking and Resource Fetching

Let's follow what happens when we type "www.example.com" into our browser and press Enter. The first major task is fetching the resources needed to display the page.

DNS Resolution

Before anything else, the browser needs to translate the human-readable domain name into an IP address that computers can use. It(network component) queries DNS (Domain Name System) servers, which act like a phone book for the internet returning an IP address like 93.184.216.34.

Establishing a Connection

With the IP address in hand, the networking component establishes a TCP connection with the server. If the URL uses HTTPS (which most modern sites do), the networking component also performs a TLS handshake to create an encrypted connection.

Sending the http request

The networking component sends an HTTP request to the server, asking for the HTML document. This request includes headers with information about our browser, what types of content it can accept, and any cookies associated with that domain. The browser engine coordinates this process, telling the networking component what to request based on the URL we entered in the user interface. The data persistence layer provides any stored cookies that need to be sent with the request.

Receiving the Response

The server responds with the HTML document, along with response headers that might include caching instructions, content type, and more. The networking component receives this data and passes it to the browser engine, which then hands it off to the rendering engine for processing.

Fetching Additional Resources

As the rendering engine parses the HTML, it discovers references to other resources: CSS stylesheets, JavaScript files, images, fonts, and more. When it encounters a <link>, <script>, or <img> tag, it notifies the browser engine, which instructs the networking component to begin fetching these resources, often in parallel to speed things up

Building the DOM : HTML Parsing

Once the browser has the HTML document, the rendering engine begins parsing it to create the Document Object Model (DOM). But what does parsing actually mean?

Understanding Parsing: A Simple Example

To understand parsing, let's use a simpler example than HTML. Imagine we need to evaluate the mathematical expression: 3 + 4 * 2

A parser's job is to understand the structure and meaning of this expression. It needs to know that:

  • There are three numbers: 3, 4, and 2

  • There are two operators: + and *

  • Multiplication has higher precedence than addition

  • The correct evaluation is 3 + (4 2) = 11, not (3 + 4) 2 = 14

The parser breaks down the text into meaningful pieces (tokens): 3, +, 4, *, 2. Then it builds a tree structure that represents the correct order of operations:

   `+
  / \\
 3   *
    / \\
   4   2`

This tree clearly shows that we multiply 4 and 2 first, then add 3 to the result.

HTML Parsing Works Similarly

HTML parsing follows a similar process, but with more complex rules. The browser reads the HTML text character by character, identifying tags, attributes, and content. It breaks the HTML into tokens (start tags, end tags, text content, comments, etc.) and builds a tree structure called the DOM.

For example, this simple HTML:

html

<html> <body> <h1>Hello World</h1> <p>This is a paragraph.</p> </body> </html>

Gets parsed into a tree structure:

Document └── html └── body ├── h1 │ └── "Hello World" └── p └── "This is a paragraph."

The DOM Tree

The DOM (Document Object Model) is a tree-like representation of the HTML document. Each element becomes a node in this tree, with parent-child relationships that mirror the nesting in the HTML. The DOM is more than just a data structure—it's a live representation of the page that JavaScript can manipulate. When JavaScript changes the DOM, those changes are reflected on the screen.

The HTML parser is remarkably forgiving. Unlike strict programming languages, HTML parsing is designed to handle errors gracefully. If we forget to close a tag or nest elements incorrectly, the browser will make its best guess about what we intended. This is why broken HTML often still displays something, even if it's not quite what we wanted.

Styling with CSS: Building the CSSOM

While the HTML parser is building the DOM, another process is happening in parallel: CSS parsing and the creation of the CSSOM (CSS Object Model).

Fetching and Parsing CSS

When the HTML parser encounters a <link> tag pointing to a CSS file, or a <style> tag containing CSS rules, the browser fetches (if external) and parses that CSS. CSS parsing works similarly to HTML parsing: the browser reads the CSS text, breaks it into tokens (selectors, properties, values), and builds a tree structure.

The CSSOM Tree

The CSSOM is similar to the DOM but represents the styling information. It's a tree structure where each node represents a CSS rule that applies to elements in the DOM. The CSSOM includes not just the CSS we wrote, but also the browser's default styles (user agent stylesheet) and any inline styles.

Why CSS Parsing Blocks Rendering

CSS parsing is render-blocking, meaning the browser won't start rendering the page until it has processed all the CSS. Why? Because if the browser started painting the page without all the styles, it might display unstyled content that then suddenly changes when the CSS loads, creating a jarring flash of unstyled content. By waiting for the CSS, the browser can render the correctly styled page from the beginning.

The Render Tree: Bringing DOM and CSSOM Together

Now we have two trees: the DOM (representing structure and content) and the CSSOM (representing styles). The next step is combining them into the render tree.

Constructing the Render Tree

The render tree is a visual representation of the document. It contains only the nodes that will be visible on the page, with their styling information attached. Here's what happens:

  1. The browser traverses the DOM tree from the root

  2. For each visible node, it finds the matching CSS rules from the CSSOM

  3. It combines the DOM node with its computed styles

  4. It creates a render tree node with this information

What's Included and What's Not

Not everything from the DOM makes it into the render tree:

  • Elements with display: none are excluded (they take up no space)

  • Elements in the <head> section are excluded (they're not visual)

  • Hidden elements like <script> and <meta> tags are excluded

However, elements with visibility: hidden ARE included in the render tree because they still take up space on the page—they're just invisible.

Each node in the render tree contains content and computed styles. For text, it knows the font, size, color, and content. For a box, it knows dimensions, colors, borders, and more.

Layout: Calculating Positions and Sizes

Having the render tree tells us what should be displayed and how it should look, but it doesn't yet tell us where things should be positioned or how large they should be. That's the job of the layout process, also called reflow.

The Layout Process

Layout is the process of calculating the exact position and size of each element. The browser starts at the root of the render tree and traverses it, calculating:

  • The exact size of each element

  • The exact position of each element

  • How elements affect each other's positions

The Box Model

Every element is treated as a rectangular box with:

  • Content: the actual content (text, images, etc.)

  • Padding: space inside the element, around the content

  • Border: a line around the padding

  • Margin: space outside the element, separating it from other elements

Layout calculates these dimensions for every element, accounting for things like:

  • Parent container sizes

  • Sibling elements

  • Font sizes and text content

  • CSS rules like min-width, max-width, width, height

Coordinate System

The layout process produces exact coordinates for each element. A <div> might end up at position (100px, 200px) with a width of 300px and height of 150px. These coordinates form a coordinate system where (0, 0) is typically the top-left corner of the viewport.

Reflow Can Be Expensive

Layout is computationally expensive, especially for complex pages. That's why browsers try to be smart about when they recalculate layout. If we change something that doesn't affect layout (like a color), the browser doesn't need to reflow. But if we change dimensions, positions, or do something that affects the layout of other elements, a reflow might be necessary. This is why excessive JavaScript manipulation of element dimensions can slow down a page—it triggers many reflows.

Painting: Putting Pixels on Screen

After layout, the browser knows exactly what should be displayed, how it should look, and where it should be positioned. Now it's time to actually paint pixels on the screen.

The Painting Process

Painting is the process of filling in pixels. The browser walks through the render tree and converts each node into actual pixels on the screen. This involves:

  1. Creating layers: The browser might create multiple layers for different parts of the page (for example, fixed position elements might be on a separate layer)

  2. Paint operations: For each layer, the browser performs paint operations:

    • Drawing backgrounds

    • Drawing borders

    • Drawing text

    • Drawing images

    • Drawing shadows and effects

  3. Rasterization: The browser converts these vector operations into actual pixels (raster graphics). This is where abstract shapes become concrete colored squares on our screen.

Painting Order

Elements don't just appear in random order. The browser follows a specific painting order (often called the stacking context):

  1. Background colors

  2. Background images

  3. Borders

  4. Children elements

  5. Outlines

Elements with higher z-index values are painted later, appearing on top.

Compositing

After painting each layer, the browser needs to combine them in the correct order to create the final image we see. This is called compositing. The browser's compositor takes all the painted layers and draws them in the right order onto the screen.

The Display Step

Finally, the composited image is sent to our screen. The pixels that the browser has been calculating are now actually lighting up physical pixels on our monitor. This is the culmination of all the previous steps—our browser has successfully transformed HTML, CSS, and JavaScript into a visual webpage.

Conclusion

A web browser is a marvel of modern software engineering. In the fraction of a second between typing a URL and seeing a page, our browser has performed DNS lookups, established secure connections, downloaded resources, parsed thousands of lines of code, built multiple tree structures, calculated the position and size of hundreds of elements, and painted millions of pixels—all while managing security, executing JavaScript, and keeping our data private. The browser is more than just a window to the internet—it's a sophisticated platform that has fundamentally changed how we communicate, work, and access information in the modern world