Raw Mode
Performance

Fast Node.js is mostly about doing less

Ishtmeet Singh//32 min read
#nodejs#performance#caching

Making a small Node.js feature fast by asking what the code actually does.

You might have pasted a link into a Slack or Discord channel, and a little card shows up under it with the page title, a short description and an image. People call it an unfurl, or (usually) a link preview. The reason for this entire blog post is that I built it for one of my (commercial) projects, and this feature turned out to be one of the most fun things I've ever built!

It looks like nothing, but it's really interesting because the first version would prolly take ten minutes, and then almost every line of it turns out to be wrong in a cool way once real people start using it. In this post I'm going to build it the same way I implemented it in that project. I'll write the version I wrote as my first implementation, then come back to it again and again, fixing one thing each time and explaining why. By the end we'll have something I'd be happy to ship, and it still fits in one small file.

A couple of rules, I'm not going to necessarily talk about using a framework or a library. I've condensed the entire feature into one function, and the HTTP part is a couple of lines you could drop into Express, Fastify, Hono or the framework your co-worker built over a weekend and swears is "basically Express, but better". Also, to be clear I'm not going to talk about measuring performance. That might be a post in the future. Every change we'll make will come from reading the code and asking a very simple question - what does this actually do when a lot of people use it at once?

You'll need Node 20 or newer, since I'm using the built-in fetch, and one dependency, cheerio, for reading HTML.

What we're building

The whole feature is one function -

js
const preview = await getLinkPreview('https://thenodebook.com/some-article');
// {
//   url: 'https://thenodebook.com/some-article',
//   title: 'Some Article',
//   description: 'A short summary of the article.',
//   image: 'https://thenodebook.com/cover.jpg',
//   siteName: 'NodeBook'
// }

Where does that data come from? Most sites put it in <meta> tags inside the <head> of the page, using the Open Graph format (that's what "og" stands for) Facebook introduced years ago -

html
<head>
  <title>Some Article | NodeBook</title>
  <meta property="og:title" content="Some Article">
  <meta property="og:description" content="A short summary of the article.">
  <meta property="og:image" content="/cover.jpg">
  <meta property="og:site_name" content="NodeBook">
</head>

So the job is - download the page, search for those tags, and return what's in them. If a page has no Open Graph tags, we'll fall back to the regular <title> and <meta name="description">.

In our app, the client calls an endpoint when you paste a link into the message box, so you can see the card before you hit send. After the message is posted, everyone in the channel sees it, and each of their clients asks the same endpoint for the preview so it can draw the card.

Version 1 - the one I wrote in a couple of minutes

I always start with the part that won't change ever, that is pulling data out of the HTML. I'll write it once and not touch it again for the rest of the post. Every improvement we make will happen around this piece of code.

js
// extract.js
export function extract($, pageUrl) {
  const meta = (key) =>
    $(`meta[property="${key}"]`).attr('content') ??
    $(`meta[name="${key}"]`).attr('content');

  return {
    url: pageUrl,
    title: meta('og:title') ?? ($('title').first().text().trim() || null),
    description: meta('og:description') ?? meta('description') ?? null,
    image: absolute(meta('og:image') ?? meta('twitter:image'), pageUrl),
    siteName: meta('og:site_name') ?? new URL(pageUrl).hostname,
  };
}

function absolute(href, base) {
  try {
    return href ? new URL(href, base).href : null;
  } catch {
    return null;
  }
}

I'm not using jQuery here. $ is a cheerio document, which gives us jQuery-style selectors over parsed HTML. The meta helper checks both property= and name= because sites are inconsistent about which one they use. The absolute helper is there because a lot of sites write og:image as a relative path like /cover.jpg, which is useless to a client sitting on a different domain.

Now let's write the feature itself

js
// different file, maybe preview.js?
import * as cheerio from 'cheerio';
import { extract } from './extract.js';

export async function getLinkPreview(url) {
  const res = await fetch(url);
  const html = await res.text();
  const $ = cheerio.load(html);
  return extract($, res.url);
}

And here's the server endpoint -

js
// server.js
import { createServer } from 'node:http';
import { getLinkPreview } from './preview.js';

createServer(async (req, res) => {
  const target = new URL(req.url, 'http://localhost').searchParams.get('url');
  try {
    // the function we wrote get's called here
    const preview = await getLinkPreview(target);
    res.writeHead(200, { 'content-type': 'application/json' });
    res.end(JSON.stringify(preview));
  } catch {
    res.writeHead(400).end();
  }
}).listen(3000);

Notice I pass res.url to extract, not the URL we were given. fetch follows redirects, and res.url is where we ended up. If someone pastes a bit.ly link, relative image paths should be resolved against the real page, not the shortener.

Ok so this works now. Paste a link, get a card. And to be honest, for a side project with ten users, I'd ship this and move on. Nothing here is stupid. It's short, it reads well, and it's easy to change. So let's go through what happens when it meets a REAL channel.

Version 2 - stop doing the same work twice

Let's say there's a channel/group with 300 people in it. Someone posts a link to a news article. Here's what happens -

  1. The person who posted it already fetched the preview once while typing.
  2. The message goes out to 300 clients.
  3. Each of those 300 clients calls /preview?url=....
  4. Our server downloads the same article 300 times, parses it 300 times, and sends back the same JSON 300 times.

That's just ONE message. Popular links, like a news story or a GitHub release, get posted in many channels. The work we do for the second request is exactly the work we did for the first, with exactly the same result. That's the first thing I look for when trying to optimise for performance i.e work that produces the same answer over and over.

The obvious fix is a cache. The simplest one might look something like this -

js
const cache = new Map();

export async function getLinkPreview(url) {
  if (cache.has(url)) return cache.get(url);

  const res = await fetch(url);
  const html = await res.text();
  const preview = extract(cheerio.load(html), res.url);

  cache.set(url, preview);
  return preview;
}

This is a big improvement, and also a slow-motion bug. Three things are wrong with it.

It never forgets anything. A Map holds on to every key you put in it until the process dies. Every unique link anyone ever pastes stays in memory. A busy app sees a lot of unique links, so this Map grows for as long as the server is up. That's just going to end up in a memory leak, just a tidy one.

Secondly, nothing ever expires. So, if an article's title gets corrected an hour after publishing, we'll keep showing the typo until the next deploy.

And last, the keys are sloppy. https://TheNodeBook.com/post, https://thenodebook.com/post and https://thenodebook.com/post#comments are the same page, but they're three different strings, so they're three cache entries and goint to end up in three different fetches.

The keys are the easiest to fix. The URL class already does most of the work i.e it lowercases the hostname for us, and we can drop the fragment ourselves, since the part after # is never sent to the server anyway.

js
function normalize(raw) {
  const url = new URL(raw);
  if (url.protocol !== 'http:' && url.protocol !== 'https:') {
    throw new TypeError('Only http and https links are supported');
  }
  url.hash = '';
  return url.href;
}

This gives us input validation for free as well. new URL(null) throws, new URL('not a url') throws, and javascript: links get rejected. Our handler already turns a throw into a 400.

Now memory. What I want is a cache with two limits - a maximum number of entries, and a maximum age for each one of the entry. When it's full, it should throw out the entry nobody has asked for in the longest time. That's an LRU cache in a nutshell, for "Least Recently Used".

You can install one from npm, and in a real project I'll obviously do taht, but it's worth seeing how small one is, because JavaScript's Map has a property that makes it almost free. A map remembers the order keys were inserted in. When you iterate over a Map, it always goes from the oldest insert to the newest. So if every time we read an entry and we delete it and put it back, the entries people use go back to the end, and the first key in the Map is always the one nobody has touched for the longest.

js
class LRU {
  constructor(max) {
    this.max = max;
    this.map = new Map();
  }

  get(key) {
    const entry = this.map.get(key);
    if (!entry) return undefined;
    if (entry.expires < Date.now()) {
      this.map.delete(key);
      return undefined;
    }
    // Move it to the back of the line.
    this.map.delete(key);
    this.map.set(key, entry);
    return entry.value;
  }

  set(key, value, ttl) {
    this.map.delete(key);
    this.map.set(key, { value, expires: Date.now() + ttl });
    if (this.map.size > this.max) {
      // The first key is the least recently used one.
      this.map.delete(this.map.keys().next().value);
    }
  }
}

That's (mostly, but not exactly) all of it. Every operation takes the same time no matter how big the cache is. There's no linked list to maintain and no timer sweeping in the background. Expired entries get cleaned up lazily. Either someone asks for them and we notice they're stale, or they drift to the front and get evicted. I like this kind of solution a lot. Instead of adding machinery, we're leaning on something the language already guarantees.

this.map.keys().next().value might look odd if you haven't seen it before. keys() returns an iterator, and next() gives us its first item without building an array. On a Map with 5,000 entries, [...this.map.keys()][0] would copy 5,000 keys into a new array just to read one of them.

With that in place -

js
const cache = new LRU(5000);
const ONE_HOUR = 60 * 60 * 1000;

export async function getLinkPreview(rawUrl) {
  const url = normalize(rawUrl);

  const cached = cache.get(url);
  if (cached) return cached;

  const res = await fetch(url);
  const html = await res.text();
  const preview = extract(cheerio.load(html), res.url);

  cache.set(url, preview, ONE_HOUR);
  return preview;
}

What's the most memory this can use? A preview is a bunch of short strings, maybe a kilobyte? Five thousand of them is a few MBs. That's a number I can reason about, and it doesn't grow no matter how long the server runs. Whenever you add a cache, you should be able to answer "what's the most memory this can ever hold?" If you can't, you're probably building a leak with extra steps.

Back to our channel of 300 people. The first client fetches the article and the other 299 get it from memory.

Or do they?

Version 3 - cache the promise, not the value

Let's try to trace exactly what happens when 300 requests arrive at almost the same moment. This is where knowing how Node runs your code really pays off. If you've been following NodeBook, you'd already know most of the below.

Node runs your JavaScript on a single thread (there are worker threads too, but they don't matter here). It never runs two pieces of your code at the same time. When a function hits an await on something that isn't ready yet, it pauses, and Node goes off and runs something else maybe another request's handler, a timer, or whatever is waiting. When the thing you awaited is ready, your function picks up where it stopped.

Now look at our function with that in mind -

js
const cached = cache.get(url);      // (1) check the cache
if (cached) return cached;

const res = await fetch(url);       // (2) pause here for half a second
const html = await res.text();      // (3) and here
// ...
cache.set(url, preview, ONE_HOUR);  // (4) finally fill the cache

Request 1 runs step (1) and ther'es a cache miss. It starts the fetch at step (2) and pauses. Node picks up a new request from another client, i.e Request 2. Request 2 runs step (1), and the cache is still empty, because request 1 won't reach step (4) for another half second. So request 2 starts its own fetch and pauses. Request 3 does the same. And so on. Can you guess what's wrong here, and how would you fix this?

Nearly all 300 requests land inside that half-second window, because they're all triggered by the same message appearing on 300 screens. So nearly all of them miss the cache, and we download the article almost 300 times. We're right back at Version 1. The cache only helps requests that arrive after the first one finishes, and in a chat app, the initial burst is the whole point.

This has a name, a cache stampede, and it's one of the most common ways a cache fails in production. It works perfectly in testing, because in testing you click one link at a time.

The fix is one of my favourite ideas in Node, and you're up for a surprise (seriously!). The problem is the gap between "I've decided to fetch this" and "the result is in the cache". During that gap, nobody else can see that someone is already on it. So, what do we do? We should close that gap! Put something in the cache the moment we decide to fetch, before any await. And that something is the promise.

js
export function getLinkPreview(rawUrl) {
  const url = normalize(rawUrl);

  const cached = cache.get(url);
  if (cached) return cached;

  const promise = fetchPreview(url);
  cache.set(url, promise, ONE_HOUR);
  return promise;
}

async function fetchPreview(url) {
  const res = await fetch(url);
  const html = await res.text();
  return extract(cheerio.load(html), res.url);
}

Let's trace it again. Request 1 misses and calls fetchPreview, which starts the fetch and immediately hands back a pending promise. Request 1 stores that promise in the cache and returns it. There's no await anywhere in that sequence, so nothing else can run in the middle. Then request 2 comes in, checks the cache, and finds the promise. It returns that same promise. Its handler does await getLinkPreview(...) and ends up waiting on the very same fetch as request 1. So do the other 298.

One download. One parse. 300 answers. A few things about this change are easy to miss.

getLinkPreview isn't async anymore. It doesn't need to be TBH. It returns a promise either way, and the handler still awaits it. I dropped the keyword on purpose, to make it obvious that nothing in this function ever pauses. If someone later adds an await between cache.get and cache.set, the stampede comes back again. That stretch from get to set has to run in one go, and I'd probably leave a comment saying so.

Awaiting a promise that has already resolved is fine. An hour later, request 5,000 finds the same promise in the cache. It resolved long ago, and awaiting it just hands back the value. There's no penalty for storing promises instead of plain values.

Everyone now shares one object. All 300 handlers get the same preview object, not copies of it. If some code further down modifies it, say by shortening the title for mobile, it changes it for everyone, for the next hour. So I freeze it. Modules run in strict mode, so any code that tries to change a frozen object throws right away instead of quietly corrupting the cache.

Then there's something that's actually a bit hard. What happens when the fetch fails?

With a cached promise, a failure gets cached too. If the site was having a bad minute when the first request came in, the rejected promise sits in the cache for an hour, and everyone gets an error for an hour. Not great, isn't it?

The obvious fix is to delete the entry when the promise rejects. But think about what that does for a site that's actually down. Every request finds nothing in the cache, tries again, fails again. We're back to hammering a server that's already struggling, and every one of our users waits for a failure.

What I want is in between: remember failures, but only briefly. So when a fetch fails, I swap the entry for an empty preview with a short lifetime:

js
const OK_TTL = 60 * 60 * 1000; // 1 hour
const FAILURE_TTL = 60 * 1000; // 1 minute

export function getLinkPreview(rawUrl) {
  const url = normalize(rawUrl);

  const cached = cache.get(url);
  if (cached) return cached;

  // No await between get() and set(), or concurrent requests will all miss.
  const promise = fetchPreview(url)
    .then(Object.freeze)
    .catch(() => {
      cache.set(url, promise, FAILURE_TTL);
      return emptyPreview(url);
    });

  cache.set(url, promise, OK_TTL);
  return promise;
}

function emptyPreview(url) {
  return Object.freeze({
    url,
    title: null,
    description: null,
    image: null,
    siteName: new URL(url).hostname,
  });
}

The catch handler refers to promise, the same variable it's being assigned to. That looks like it shouldn't work, but it does. The handler can't run until the fetch has failed, which is always later, and by then promise has its value. We store that same promise again, which will now resolve to the empty preview, with a one-minute lifetime instead of an hour.

There's a product decision hiding in here too. getLinkPreview never rejects anymore. A preview is a nice extra on a message, not the message itself. If we can't get one, the client should just not draw a card. Returning an empty preview says "we tried, there's nothing to show" in a way the client can handle without an error path. Only a bad URL still throws, and that really is a 400.

Version 4 - don't wait forever

Here's a question we might never think about while writing the version 1 - how long will fetch wait for a server that isn't answering?

In Node, the answer is a long time. The built-in fetch is powered by a library called undici, and it has no overall deadline for a request. There are a few internal timeouts, but the ones that are important for the discussion are measured in minutes, not seconds. And the one that watches the body only measures the gap between chunks, so a server that sends one byte every couple of minutes can keep a request alive pretty much forever. If someone pastes a link to a server that accepts the connection and then just sits there, our request sits there with it.

In Version 1, that was one unlucky user staring at a spinner. After Version 3 it's much worse, and I want you to think about this for a moment. When you share one request between many callers, you also share its problems. Everyone who asks for that link is now awaiting the same stuck promise, so the whole channel waits, all 300 of them. And because that promise never settles, it never fails either, which means the one-minute failure cache we just built never kicks in. Every one of those client connections stays open, holding memory, for minutes.

Think about it from the user's side too. A preview that shows up after 30 seconds is useless. By then the person has read the message, replied, and moved on. So the right amount of time to wait is short. I'll give it four seconds -

js
const TIMEOUT = 4000;

async function fetchPreview(url) {
  const res = await fetch(url, {
    signal: AbortSignal.timeout(TIMEOUT),
    headers: { 'user-agent': 'ChatPreviewBot/1.0', accept: 'text/html' },
  });
  const html = await res.text();
  return extract(cheerio.load(html), res.url);
}

AbortSignal.timeout(ms) gives you a signal that aborts itself after that many milliseconds. It's built into Node, and there's no timer for us to clear. It also covers the whole request, including reading the body. This last part is more important than it looks. A server can send its headers right away and then send body chunks out a byte at a time. A timeout that only covered "time until the response starts" would never catch that, and we'd be back to waiting forever.

When it fires, fetch (or res.text()) rejects, the catch we wrote in Version 3 turns it into an empty preview, and the failure gets remembered for a minute. I really like how these two go hand-in-hand together - the timeout turns "hangs forever" into "fails quickly", and the failure cache turns "fails quickly" into "fails once".

While I was in there, I also added a user-agent. Some sites serve an empty shell or an error page to requests that don't say who they are, and a few only include their Open Graph tags for things that look like crawlers (which is exactly what we are...). It's not a speed change at all, but it stops us from doing all this work just to get nothing useful back.

Version 5 - read only what you need

This is the one thing that taught me something new the first time I built this feature. Look at these two lines. They've been sitting there, untouched, since Version 1 -

js
const html = await res.text();
return extract(cheerio.load(html), res.url);

What does res.text() actually do? It reads the entire response body into memory and hands it to you as one big string. The entire body, whatever that body happens to be.

And what do people paste into chat? Articles, sure. But also a link to a 2 GB Linux ISO. A 400 MB video. A PDF. A single-page app whose HTML is 3 MB of inlined JavaScript. A log file that somebody's server is exposing by accident. People paste everything.

For that ISO, res.text() will happily try to read two gigabytes into a string. The four-second timeout from Version 4 helps, but on a fast connection four seconds is plenty of time to pull hundreds of megabytes into memory, for a file that has no <head> and never will. Get a few of those at the same time and the process runs out of memory and crashes, taking every other user's request down with it.

There's another problem on the parsing side as well. Say the page is real HTML, just big, like a long article with thousands of comments rendered into it, 5 MB of markup. cheerio.load(html) parses all of it into a tree of objects, and that parsing is plain synchronous JavaScript. Remember what we talked about in Version 3? Node runs your code on one thread. So while cheerio is chewing through 5 MB of HTML, nothing else in the process runs. Not the other preview requests, not even the ones that would be instant cache hits, also not your health checks too. Everybody on this server waits for one big page to finish parsing.

This is the thing about blocking the event loop that I talk a lot about in NodeBook. It doesn't just make that request slow. It makes every request slow, including the ones that have nothing to do with it and all of that work is wasted. Everything we need lives inside <head>, which is usually the first few kilobytes of the page. We're downloading and parsing an entire document just to read its first paragraph.

So let's stop doing that and be a "100x" engineer (/s). There are two changes.

First, let's look at what we got before reading it. The Content-Type header arrives before the body does. If it isn't HTML, there's no <head> to find, so there's no reason to read anything. And if it's an image, we get a nicer preview for free, because an image is its own preview.

js
const type = res.headers.get('content-type') ?? '';
if (!type.includes('text/html')) {
  await res.body?.cancel();
  return type.startsWith('image/')
    ? { ...emptyPreview(res.url), image: res.url }
    : emptyPreview(res.url);
}

That res.body?.cancel() is important, and it's a Node-specific thing that you need to know about. When fetch gives you a response, the body hasn't been downloaded yet. It's a stream, sitting there waiting for you to read it. If you just walk away from it, Node's fetch doesn't close the connection right away. The socket stays tied up until the garbage collector eventually gets around to that response object, which can take a while. Under load, those half-used connections pile up. So the rule is - if you're not going to read a body, cancel it. That tells Node to drop the connection now.

Second, for HTML, read only until </head>. res.body is a stream of chunks, handed to us as they come off the network. Instead of waiting for all of them, we can read one at a time, look for the end of the head, and stop the moment we see it.

js
const MAX_HEAD_BYTES = 512 * 1024;

async function readHead(body, limit) {
  const reader = body.getReader();
  const decoder = new TextDecoder();
  let html = '';
  let bytes = 0;

  try {
    while (bytes < limit) {
      const { done, value } = await reader.read();
      if (done) break;

      bytes += value.byteLength;
      const searchFrom = Math.max(0, html.length - 6);
      html += decoder.decode(value, { stream: true });

      const end = html.slice(searchFrom).search(/<\/head|<body/i);
      if (end !== -1) return html.slice(0, searchFrom + end);
    }
    return html;
  } finally {
    reader.cancel().catch(() => {});
  }
}

There's a lot more going on in these twenty lines than it looks like, so let's go through it slowly.

TextDecoder turns bytes into text. The { stream: true } option is there because of how UTF-8 works. Any character outside plain ASCII (an accented letter, or maybe an emoji) takes more than one byte, and a network chunk can end right in the middle of one. Without stream: true, a character split across two chunks turns into garbage. With it, the decoder holds on to the leftover bytes and finishes the character when the next chunk arrives.

Now onto the the search part. The obvious way to write it is to run html.search(/<\/head/i) on the whole string after every chunk. That works, but it re-scans everything we've already checked, every single time. With lots of small chunks, you end up scanning the start of the page over and over again. We only need to look at the new text. But there's an interesting thing that you might have already thought about. The tag itself can be split between two chunks, i.e one chunk ends with </he and the next one starts with ad>. If we only searched the new text, we'd miss it completely. So we start searching a few characters before the new text, just enough to catch a tag that straddles the boundary. </head is six characters long, so going back six is enough.

The regex also matches <body, because plenty of real pages never close their head at all, and once the body starts, the head is over either way.

limit is kind of a safety measure. Some sites stuff hundreds of kilobytes of inline scripts and styles into their <head>. I'd rather stop and parse what we have than keep reading forever. 512 KB is generous, since most pages put their meta tags in the first few kilobytes.

And finally { reader.cancel() } is the same lesson as before. Once we're done, whether we found the end of the head, hit the limit, or got an error, we tell Node we don't want the rest. For a 5 MB page, we read maybe 20 KB and hang up, instead of downloading the other 4.98 MB.

Here's what fetchPreview looks like now -

js
async function fetchPreview(url) {
  const res = await fetch(url, {
    signal: AbortSignal.timeout(TIMEOUT),
    headers: { 'user-agent': 'ChatPreviewBot/1.0', accept: 'text/html' },
  });

  if (!res.ok) {
    await res.body?.cancel();
    throw new Error(`Got ${res.status} from ${url}`);
  }

  const type = res.headers.get('content-type') ?? '';
  if (!type.includes('text/html')) {
    await res.body?.cancel();
    return type.startsWith('image/')
      ? { ...emptyPreview(res.url), image: res.url }
      : emptyPreview(res.url);
  }

  const head = await readHead(res.body, MAX_HEAD_BYTES);
  return extract(cheerio.load(head), res.url);
}

I also added in a res.ok check too, with the same cancel. A 404 page's title is "Page not found", and nobody wants to see that in a card.

Now look at what changed for the parser. We're still using cheerio, and extract hasn't changed at all (I told you we wouldn't touch it again). We're just handing it the head of the page instead of the whole thing. Cheerio doesn't mind a document that stops partway through. HTML parsers are built to cope with broken markup, because the web is full of it. The parse that used to freeze the whole process for the length of a 5 MB document now runs over a few kilobytes. We didn't make the parser perform any faster, we gave it less to do.

I think that's the most useful idea in this entire post. When something is slow, the first instinct is to find a faster way to do the same work, maybe a faster library or a smarter algorithm. But more often, the better question is whether you need to do that work at all. We didn't need the body. We didn't need the ISO. And we didn't need 299 of the 300 fetches.

Where it ended up

Here's the whole thing in one file -

js
// preview.js
import * as cheerio from 'cheerio';

const OK_TTL = 60 * 60 * 1000; // 1 hour
const FAILURE_TTL = 60 * 1000; // 1 minute
const TIMEOUT = 4000;
const MAX_HEAD_BYTES = 512 * 1024;

class LRU {
  constructor(max) {
    this.max = max;
    this.map = new Map();
  }

  get(key) {
    const entry = this.map.get(key);
    if (!entry) return undefined;
    if (entry.expires < Date.now()) {
      this.map.delete(key);
      return undefined;
    }
    this.map.delete(key);
    this.map.set(key, entry);
    return entry.value;
  }

  set(key, value, ttl) {
    this.map.delete(key);
    this.map.set(key, { value, expires: Date.now() + ttl });
    if (this.map.size > this.max) {
      this.map.delete(this.map.keys().next().value);
    }
  }
}

const cache = new LRU(5000);

export function getLinkPreview(rawUrl) {
  const url = normalize(rawUrl);

  const cached = cache.get(url);
  if (cached) return cached;

  // No await between get() and set(), or concurrent requests will all miss.
  const promise = fetchPreview(url)
    .then(Object.freeze)
    .catch(() => {
      cache.set(url, promise, FAILURE_TTL);
      return emptyPreview(url);
    });

  cache.set(url, promise, OK_TTL);
  return promise;
}

async function fetchPreview(url) {
  const res = await fetch(url, {
    signal: AbortSignal.timeout(TIMEOUT),
    headers: { 'user-agent': 'ChatPreviewBot/1.0', accept: 'text/html' },
  });

  if (!res.ok) {
    await res.body?.cancel();
    throw new Error(`Got ${res.status} from ${url}`);
  }

  const type = res.headers.get('content-type') ?? '';
  if (!type.includes('text/html')) {
    await res.body?.cancel();
    return type.startsWith('image/')
      ? { ...emptyPreview(res.url), image: res.url }
      : emptyPreview(res.url);
  }

  const head = await readHead(res.body, MAX_HEAD_BYTES);
  return extract(cheerio.load(head), res.url);
}

async function readHead(body, limit) {
  const reader = body.getReader();
  const decoder = new TextDecoder();
  let html = '';
  let bytes = 0;

  try {
    while (bytes < limit) {
      const { done, value } = await reader.read();
      if (done) break;

      bytes += value.byteLength;
      const searchFrom = Math.max(0, html.length - 6);
      html += decoder.decode(value, { stream: true });

      const end = html.slice(searchFrom).search(/<\/head|<body/i);
      if (end !== -1) return html.slice(0, searchFrom + end);
    }
    return html;
  } finally {
    reader.cancel().catch(() => {});
  }
}

function extract($, pageUrl) {
  const meta = (key) =>
    $(`meta[property="${key}"]`).attr('content') ??
    $(`meta[name="${key}"]`).attr('content');

  return {
    url: pageUrl,
    title: meta('og:title') ?? ($('title').first().text().trim() || null),
    description: meta('og:description') ?? meta('description') ?? null,
    image: absolute(meta('og:image') ?? meta('twitter:image'), pageUrl),
    siteName: meta('og:site_name') ?? new URL(pageUrl).hostname,
  };
}

function absolute(href, base) {
  try {
    return href ? new URL(href, base).href : null;
  } catch {
    return null;
  }
}

function normalize(raw) {
  const url = new URL(raw);
  if (url.protocol !== 'http:' && url.protocol !== 'https:') {
    throw new TypeError('Only http and https links are supported');
  }
  url.hash = '';
  return url.href;
}

function emptyPreview(url) {
  return Object.freeze({
    url,
    title: null,
    description: null,
    image: null,
    siteName: new URL(url).hostname,
  });
}

And the server file? It didn't change at all.

Let's go back to our channel of 300 people one last time, and trace again what happens now when that link lands in it -

  1. The first request normalizes the URL, misses the cache, and puts a promise there straight away.
  2. The other 299 find that promise and wait on it.
  3. One fetch goes out, with a four-second limit.
  4. If the response is HTML, we read until </head>, usually a few kilobytes, and hang up.
  5. Cheerio parses those few kilobytes in no time, so the event loop stays free for everyone else.
  6. All 300 requests get the same frozen object.
  7. For the next hour, anyone who posts or sees that link gets it straight from memory.
  8. If anything went wrong along the way, everyone gets an empty preview, and we try again in a minute.

And the cases that used to hurt - A link to a huge file now costs us one set of headers and a cancelled connection. A server that never answers costs us four seconds, at most once a minute. The cache can't grow past 5,000 entries, no matter how long the process runs.

We started with five lines and ended up with about 150. That's a lot more code, sure. But I can point at almost any line in there and tell you exactly which real-world problem it's there for. That's the kind of growth I'm totally fine with.

What I left alone, on purpose

Knowing when to stop is part of the job too. Here's what I thought about and decided not to do, at least for now.

Stale-while-revalidate. When an entry expires after an hour, the next person to ask has to wait for a fresh fetch. A nicer approach is to hand them the old preview immediately and refresh it in the background. It's a good idea, and it fits really well with the promise cache, but one slow preview per link per hour isn't worth the extra code yet.

A shared cache. Our cache lives in the memory of one Node process. Run four processes behind a load balancer and you have four separate caches, so a popular link gets fetched up to four times. For a lot of apps that's completely fine. If it isn't, the next step is Redis (or valkey or something like it), and the promise trick still has its place inside each process, so you don't stampede Redis either.

Storing the preview on the message. Once a message is sent, the server could fetch the preview once and save it along with the message, so the 300 clients never have to ask at all. That's a good design, and I'd probably end up there eventually. But the endpoint is still needed for the composer (the card you see while you're still typing), and everything in this post applies to that job just the same.

A streaming parser. Right now we collect the head into a string and then parse it. A streaming HTML parser like htmlparser2 could handle chunks as they arrive, pick up meta tags as it passes them, and never build a string or a tree at all. That would be faster still. But we're only parsing a few kilobytes now, and switching parsers means rewriting extract. The big win was reading less, and we already have that.

Limiting outgoing requests. If a thousand different links got posted in the same second, we'd start a thousand fetches at once. A small concurrency limit would smooth that out. I haven't seen a chat app where that happens often enough to matter, so I'd wait until it does.

There's one thing that has nothing to do with speed, but that you must not skip. A link preview endpoint fetches any URL a user gives it, from inside your network. Sooner or later someone will paste http://169.254.169.254/ (the cloud metadata address, where credentials tend to live) or http://localhost:6379/ (hello, Redis) just to see what your server can reach. Before you ship this, resolve the hostname and refuse private and internal addresses. And remember that fetch follows redirects, so checking only the URL you were given isn't enough, an innocent-looking link can redirect straight to an internal address. This attack is called SSRF (Server-Side Request Forgery), it's a whole topic of its own, and it's the one part of this feature I'd never leave for later.

The short version

If you take only one thing from this post, take the questions, not the code. Every version came from asking one plain question about what the code was really doing -

  • Am I doing the same work more than once? Cache it.
  • What happens if lots of requests arrive before the first one finishes? Cache the promise.
  • What happens when the other side never answers? Give it a deadline.
  • Am I doing work I don't need, and what does it cost everyone else while I do it? Read less, parse less.

None of that needed a framework, a faster library or a clever algorithm. All it needed was for me to picture 300 people opening the same message at the same moment, and to read my own code slowly.