All posts

How to Convert a Web Page to Markdown

Convert any URL to markdown with Pandoc, curl, browser clippers, or a reader service — how to strip navigation and ads, fix image links, and handle JavaScript-rendered pages.

Short answer: To convert a URL to markdown, pipe the page through Pandoc: curl -sL https://example.com | pandoc -f html -t gfm --wrap=none -o page.md. Pandoc also accepts a URL directly. That gives you the whole page including navigation and footers — for just the article text, use a browser clipper extension or a reader-mode service that extracts the main content first.

Converting web pages to markdown has gone from a niche archival habit to something people do daily. You want documentation in your notes. You want a reference article in your repo. You want to hand an article to an AI tool without pasting 40KB of HTML. Here's how to do it well, and the parts that don't work.


Why You'd Want This

Feeding documents to AI tools. This is the big one now. Language models process markdown far more efficiently than HTML — the same content costs a fraction of the tokens once you strip the tags, and the structure survives. If you're building a context folder for Claude Code or a similar agent, converted markdown is the right storage format. More on that in why AI tools use markdown.

Archiving. Pages disappear. A markdown copy in your notes folder outlives the site and diffs cleanly if the page changes.

Note-taking and research. Clip the useful parts of a reference page into your notes and keep them searchable with the rest of your markdown files.

Content migration. Moving a site off a CMS onto a static site generator means converting every page to markdown. Doing it by hand is not viable past about five pages.


Method 1: Pandoc

Pandoc reads HTML and writes markdown. It's the most direct route, and it runs entirely locally.

brew install pandoc

Pandoc accepts a URL as its input source:

pandoc -f html -t gfm --wrap=none https://example.com/article -o article.md

Or fetch with curl and pipe, which gives you control over headers, redirects, and cookies for pages behind a login:

curl -sL -H "User-Agent: Mozilla/5.0" https://example.com/article \
  | pandoc -f html -t gfm --wrap=none -o article.md

The flags that matter:

  • -f html tells Pandoc the input is HTML rather than guessing from a nonexistent extension.
  • -t gfm produces GitHub Flavored Markdown. -t commonmark is stricter; -t markdown_strict produces the original 2004 Markdown dialect with no extensions.
  • --wrap=none keeps paragraphs on one line. Without it Pandoc hard-wraps at 72 columns, which makes every future edit produce a messy diff.

Pandoc also has --strip-comments for removing HTML comments, and -t gfm-raw_html if you want it to drop inline HTML it can't convert rather than passing it through.


The Problem: You Get the Whole Page

Run the Pandoc command on a real article and the output starts something like this:

[Home](/home) [Products](/products) [Blog](/blog) [Contact](/contact)

[Subscribe to our newsletter](/subscribe)

# The Article Title You Actually Wanted

Then the article, then the footer, the cookie notice, three related-post teasers, and the social links. Pandoc converts the HTML it's given, faithfully — it has no opinion about which part is the article. For a documentation page that's often acceptable. For a blog post it's mostly noise.


Method 2: Readability Extraction

The fix is to run the page through a content extractor first — the same technology behind Safari's Reader mode and Firefox's Reader View. These libraries score the DOM to find the block that looks like the main article body and discard everything else.

Mozilla's Readability (github.com/mozilla/readability) is the reference implementation and the one that powers Firefox Reader View. It's a JavaScript library, so using it directly means writing a small Node script that fetches the page, parses it into a DOM, runs Readability, and hands the cleaned HTML to a markdown converter — real setup for a one-off conversion. For occasional use, the two options below are better.


Method 3: Browser Clipper Extensions

The lowest-friction option: an extension that converts the current page to markdown and puts it on your clipboard or saves it to a folder. Most run Readability-style extraction plus HTML-to-markdown conversion in the browser, so you get clean article text in one click.

Obsidian Web Clipper is Obsidian's own extension, available for Chrome and Chromium browsers, Firefox, and Safari. It saves pages as markdown with configurable templates and frontmatter, and can copy to the clipboard rather than save to a vault. MarkDownload is a long-standing vendor-neutral alternative built on Turndown and Readability.js that downloads clipped pages straight to your filesystem.

The advantages over curl:

  • The page is already rendered, so JavaScript-loaded content is captured.
  • You're already logged in, so paywalled or account-gated pages you have access to work.
  • Extraction is built in, so you get the article rather than the chrome.

The disadvantage: it's manual and per-page. Not a migration tool.


Method 4: Reader Services

There's a category of hosted services that take a URL as a path prefix and return clean markdown. The best-known is Jina AI's Reader — prefix any URL with https://r.jina.ai/ and you get markdown back:

curl "https://r.jina.ai/https://example.com/article" -o article.md

The response includes a title line, the source URL, and the extracted markdown. It handles JavaScript-rendered pages, because the service renders them server-side before extracting — which makes it the fastest path for feeding pages to an AI pipeline, and it composes well in scripts. The tradeoffs:

  • You're sending the URL to a third party. For public pages that's usually irrelevant. For internal wikis, staging sites, or anything with a token in the query string, it is not.
  • Rate limits. Free unauthenticated use is throttled; sustained use needs an API key.
  • You depend on the service staying up. Fine for a one-off, risky as a load-bearing part of a build.

Other services in this category work the same way. Evaluate any of them on the same three axes.


Images: Relative vs Absolute URLs

This one causes more broken files than anything else. Pandoc preserves image src attributes exactly as it finds them. If the page uses a relative path:

<img src="/img/diagram.png" alt="Architecture diagram">

You get:

![Architecture diagram](/img/diagram.png)

Which points at nothing once the file leaves the website. The same applies to relative links between pages.

Three ways to handle it:

Rewrite to absolute. A quick sed pass over the output, substituting the site's origin for the leading slash:

sed -i '' 's|](/|](https://example.com/|g' article.md

The images now load from the original server — which works until the site moves or the file is renamed.

Download the images. Pull them locally and rewrite the paths to point at a local folder. This is what you want for genuine archival, since the markdown file and its images stay together.

Drop them. If you're converting for AI context, images add nothing and cost tokens. Strip them entirely.

Note that --extract-media, which works for DOCX and other container formats, doesn't help here. HTML doesn't embed its images; it references them.


JavaScript-Rendered Pages Don't Work

curl fetches the HTML the server sends. It does not execute JavaScript. For a single-page app, a React or Vue site, or anything that loads its content via a client-side API call, that HTML is often close to empty:

<div id="root"></div>

Pandoc converts that faithfully into nothing.

Symptoms: your markdown output is a few navigation links and a blank body, or a "Loading..." string, or a <noscript> message. If you see that, the page needs a real browser.

Options:

  • Use a browser clipper extension — the page is already rendered in your tab.
  • Save the page from the browser (File > Save Page As, complete) and convert the saved HTML file with Pandoc.
  • Use a headless browser — Playwright or Puppeteer can load the page, wait for content, dump the rendered DOM, and hand that to Pandoc. This is the scriptable option for bulk work, and considerably more setup.
  • Use a reader service that renders server-side.

Bulk Conversion

For a site migration, work from a list of URLs:

while read -r url; do
  slug=$(echo "$url" | sed 's|.*/||; s|[^a-zA-Z0-9-]|-|g')
  curl -sL "$url" | pandoc -f html -t gfm --wrap=none -o "out/${slug}.md"
  sleep 1
done < urls.txt

The sleep 1 is not optional politeness — it keeps you from hammering a server and getting blocked. Expect to post-process: every site has its own boilerplate, and a sed or Python pass that strips the nav block and footer from this site will do more for output quality than any generic tool.


Copyright and Scraping

Converting a page to markdown doesn't change its copyright status. A few practical lines:

  • Personal archiving and research use is generally uncontroversial. Saving a reference page to your own notes is what browsers have always done.
  • Republishing converted content is a copyright issue regardless of format. Convert freely, publish carefully.
  • Check robots.txt and the terms of service before bulk-fetching a site you don't own, and rate-limit yourself. Automated retrieval at scale is a different thing from reading a page.
  • Feeding content to an AI tool doesn't launder its origin. If you'd need permission to quote it, you'd need permission to publish output derived from it.

After Conversion

Converted markdown always needs a pass. Headings arrive at the wrong level because the page used <h1> for its logo. Tables come through as raw HTML because they used colspan. Code blocks lose their language tag because the site used CSS classes Pandoc doesn't map.

Read the file before you rely on it. OpenMark opens a converted .md and renders it immediately — Document view shows what actually landed, Markdown view shows the raw source for cleanup. Broken image links and orphaned nav lists are obvious in the rendered view and invisible in a wall of source.

If your goal is a reference library for an AI agent, markdown in the age of AI agents covers structuring the folder once the files exist. For structured API responses rather than pages, see how to convert JSON to markdown.


Download OpenMark → — $9.99, one-time, native macOS. Open your converted pages and see instantly what survived the trip from HTML.