Skip to content
Go back

Migrating 10 Years of Bookmarks

By KingPin 15 min read
Migrating 10 Years of Bookmarks
Contents

Your Bookmarks Live in Five Places and Trust None of Them

Somewhere on your machine there is a Pocket CSV from before the shutdown, a Chrome profile you stopped using in 2019, a Firefox profile you use now, a Raindrop account you opened during a productivity phase, and maybe a Pinboard export from the days when that was the cool kid’s choice. Ten years of “I’ll read this later” are spread across all of them. It’s 2 AM, you finally want one searchable pile, and every tool wants you to import one file at a time.

The plan is short. Export everything to Netscape HTML (the one format all of these tools speak), normalize URLs the same way your target tool does, let the importer upsert, and keep dead links tagged instead of deleted. A small stdlib Python pipeline does the middle part. No pip installs, no database, no framework.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/self-hosting/migrate-bookmarks-self-hosted

Still choosing a tool? Read linkding vs Hoarder vs Wallabag first. The target here is linkding (v1.47.0 as of October 2026), with a short aside on Karakeep.

Why Netscape HTML Is the Lingua Franca

The Netscape bookmark format dates from the 1990s and refuses to die. Folders are <H3> headings followed by a nested <DL>, and links are <A> tags with attributes:

merged.html
<DT><A HREF="https://sqlite.org/whentouse.html" ADD_DATE="1420000000" TAGS="database,source:pocket,sqlite" PRIVATE="1" TOREAD="1">Why SQLite</A>
<DD>An optional description goes here.

ADD_DATE is unix seconds, TAGS is comma-separated, and <DD> holds a description. Chrome, Firefox, Raindrop, Pinboard, linkding, and Karakeep all read and write it. Here is what each source gives you:

SourceHow to exportTags?Dates?Gotcha
Chromechrome://bookmarks, three-dot menu, Export bookmarksNoADD_DATE, LAST_MODIFIEDFolders are the only metadata, plus bulky icon data URIs
FirefoxCtrl+Shift+O, Import and Backup, Export Bookmarks to HTMLYes (TAGS)ADD_DATE, LAST_MODIFIEDIgnore the .jsonlz4 backups in bookmarkbackups/, HTML is easier
RaindropExport page, choose HTMLYesYesCollections become folders. CSV and TXT exist, but Raindrop’s own help recommends HTML
Pinboardpinboard.in/export, JSONYes (space-separated)ISO 8601 timeNeeds login. description is the title, extended is the notes
PocketAn old CSV, if you kept oneYes (pipe-separated)time_added in unix secondsPocket shut down July 8, 2025 and the export page closed in late 2025

Pocket deserves a note. Nobody can export from it today. If you have the CSV in a downloads folder, you are fine. If you don’t, that data is gone, and this is your reminder to export from every hosted service while it exists. The header is title,url,time_added,tags,status, and some exports add a cursor column after time_added. The script reads columns by name, so the extra one is harmless.

Chrome gives you folders and nothing else. Firefox and Raindrop give you real tags. Linkding has no folders at all, so the folder structure has to become tags, or you lose it.

The Normalizer

normalize.py reads each input by extension and content: .html as Netscape, .csv as Pocket (when the header has time_added) or Raindrop, .json as Pinboard. The HTML side uses html.parser.HTMLParser and tracks a stack of open <DL> levels. An <H3> names the folder, and the next <DL> pushes it onto the stack. Each <A> records the folders open at that moment:

normalize.py
if tag == "dl":
self.stack.append(self.pending)
self.pending = None
elif tag == "h3":
self.buf = ""
elif tag == "a":
tags = [t.strip() for t in a.get("tags", "").split(",") if t.strip()]
self.cur = rec(a.get("href", ""), add=to_unix(a.get("add_date")), tags=tags,
folders=[f for f in self.stack if f], toread=a.get("toread") == "1",
private=a.get("private") != "0", source=self.source)
self.records.append(self.cur)

Every other source gets converted into the same record shape: URL, title, add date, tag set, folders, to-read flag, private flag, description, and source name. After that the rest of the pipeline doesn’t care where a record came from.

Folder names become tags

With --folder-tags, each folder in the path turns into a tag: lowercased, spaces to hyphens. A bookmark in Home Lab / Reverse Proxies gets home-lab and reverse-proxies. The browser’s root folders (“Bookmarks bar”, “Other bookmarks”, “Bookmarks Menu” and friends) are skipped, because a tag named bookmarks-bar on 3,000 links helps nobody. Tags the bookmark already had stay.

If your folders go six levels deep, the top levels are usually noise. --min-tag-depth N skips the first N-1 levels of the path (after the root folders are removed), so --min-tag-depth 2 keeps only the second level and below.

Dedup like linkding does

Dedup only works if your key matches the importer’s key. Linkding’s normalize_url lowercases the scheme and host, strips trailing slashes from the path, and sorts query parameters alphabetically. It keeps the fragment and keeps www.. The script copies those rules:

normalize.py
def normalize_url(url):
"""Mirror linkding: lowercase scheme and host, strip trailing slashes, sort query params."""
p = urlsplit(url.strip())
query = urlencode(sorted(parse_qsl(p.query, keep_blank_values=True)), quote_via=quote)
return urlunsplit((p.scheme.lower(), p.netloc.lower(), p.path.rstrip("/"), query, p.fragment))

Why match exactly? Linkding skips a second occurrence of the same normalized URL inside one file and counts it as failed. If your script is stricter than linkding (say, it strips www.), you merge records that linkding would have kept apart. If it’s looser, the importer logs failures for the rest. Matching the key means your “duplicates merged” number lines up with what the target would have done anyway.

www.example.com and example.com stay separate bookmarks. That’s linkding’s behavior, so it’s the script’s behavior. If you want them merged, do it by hand after the import.

The merge rule

When two records share a normalized URL, the script keeps one and folds the other in:

normalize.py
old["dups"] += 1
stats["merged"] += 1
old["tags"] |= r["tags"]
old["add"] = min([d for d in (old["add"], r["add"]) if d] or [0])
old["title"] = max(old["title"].strip(), r["title"].strip(), key=len)
old["desc"] = max(old["desc"], r["desc"], key=len)
old["toread"] |= r["toread"]
old["private"] |= r["private"]

The earliest add date wins, so a link you saved in 2014 in Chrome and again in 2021 in Raindrop shows up as 2014. Tags are unioned, the longest title wins (a bare domain loses to a real page title), and the to-read and private flags stay on if either copy had them. Private is the safer default: you can make a bookmark public later, but you can’t un-publish one.

Each record also gets a source:<name> tag (source:chrome, source:pocket, and so on). After the import you can search #source:pocket and see what that old export contributed. For HTML inputs the name is the filename stem, so call your files chrome.html and firefox.html.

Some hrefs are not web pages: javascript: bookmarklets, place: queries, chrome://, about:, file://, and data:. Those are dropped and counted.

Source-specific mapping

A few fields need translation:

Running It

Terminal window
python3 normalize.py --folder-tags --out merged.html /path/to/exports/*

The summary goes to stderr. With the sample files shipped in the repo it looks like this:

read chrome: 4
read firefox: 3
read pinboard: 3
read pocket: 3
read raindrop: 3
skipped non-http: 1
duplicates merged: 1
written: 14
distinct tags: 22

Sixteen records in, one bookmarklet dropped, one duplicate merged, fourteen out. On a real decade of bookmarks the duplicate count is the interesting number. Write down the “written” figure now, because you’ll compare it to linkding’s count at the end.

The output is a flat file with one <DL>, no folders, and every bit of structure carried as tags. The repo also ships test_normalize.py, a plain assert script that runs the normalizer over the samples and checks the bookmarklet drop, the Chrome and Firefox merge, folder tags, the Pinboard to-read flag, and the Pocket cursor column. Run python3 test_normalize.py and expect ok.

Ten years is long enough for a large share of your links to go dark. You could delete the dead ones. Don’t. A dead URL is still evidence: the title and your tags tell you what the page was about, and the Wayback Machine may have a copy. Deleting is permanent and tagging costs nothing, so the pipeline tags.

linkcheck.py reads the merged file and checks each URL with a HEAD request, falling back to GET when the server answers 403, 405, or 501. A thread pool handles the volume:

linkcheck.py
if status in (403, 429):
cls = "blocked"
elif status in (404, 410):
cls = "gone"
elif status >= 500:
cls = "error"
elif 200 <= status < 300:
same = urlsplit(final).hostname == urlsplit(url).hostname
cls = "ok" if same else "moved"
else:
cls = "error"

The classes are ok, moved (the final host differs from the original, so a domain change or a parked domain), gone (404 or 410), error (5xx, DNS failure, timeout, SSL error), and blocked. A 403 or 429 is not a dead page. It’s a CDN deciding your Python script looks like a bot. Treating it as dead would tag half of Cloudflare’s customers.

Terminal window
python3 linkcheck.py merged.html --out report.csv --workers 16 --timeout 10

The CSV columns are url,status_class,http_status,final_url,wayback_url,error, and counts per class go to stderr. Sixteen workers is polite enough for most hosts. If many of your links share one site, drop the count.

The Wayback lookup

Add --wayback and the script queries https://archive.org/wayback/available for gone and error rows only. Wayback rate-limits hard. A handful of quick requests from one IP returned HTTP 429 while I was checking this. So the lookups run one at a time with a 1.5 second sleep, and a 429 is recorded as rate-limited, which means “unknown, try again later” and never “not archived”. The cost is real: 400 dead links take about ten minutes. Run it in a terminal multiplexer and go make coffee.

The lookup returns a closest snapshot URL when one exists, and an empty archived_snapshots object when nothing is archived. Snapshot links go into the wayback_url column.

Test the base run without --wayback first. Then add it once you’ve seen the class counts.

Importing Into linkding

You have two ways in.

The UI. Settings, then the Import section, upload merged.html. It works for a few hundred links. For thousands, the upload runs inside one HTTP request, and linkding’s uwsgi request timeout is LD_REQUEST_TIMEOUT, 60 seconds by default. The linkding docs say to raise it if a large file hits request timeouts. In your compose file:

environment:
- LD_REQUEST_TIMEOUT=300

The management command. This skips HTTP entirely and is what I’d use for anything over a few hundred links:

Terminal window
docker cp merged.html linkding:/tmp/merged.html
docker compose exec linkding python manage.py import_netscape /tmp/merged.html you

The two positional arguments are the file and the linkding username. No timeout, no browser tab to babysit.

What the importer does with your file

These behaviors come from linkding’s importer source (v1.47.0):

Upsert is the reason to import everything in one merged file. Import the same file twice and nothing multiplies. Fix a title in the file, import again, and the existing bookmark updates.

A Karakeep Aside

If you’d rather have AI tagging, Karakeep (v0.33.2 as of October 2026) imports Netscape HTML, Pocket’s CSV, and Omnivore JSON from Settings. Titles, tags, and add dates are preserved, and an automatically created list holds everything you imported. You can’t pick and choose, so every URL in the file goes in, dead ones included. Run the dead-link triage first and trim the file. Every import queues crawling and AI tagging, so a 5,000-link file keeps the worker busy for hours. See the Karakeep review for how that tagging behaves.

The Tagging Pass Over the API

Back to linkding. After the import, tag the dead ones. The REST API needs a token from Settings, Integrations. GET /api/bookmarks/check/?url= tells you whether a URL exists and returns the bookmark, and PATCH /api/bookmarks/<id>/ with tag_names updates the tags. Here is tag_dead_links.py in full:

tag_dead_links.py
#!/usr/bin/env python3
"""Add dead-link (and has-wayback) tags in linkding from a linkcheck.py report. Deletes nothing."""
import csv, json, os, sys
import urllib.parse, urllib.request
BASE = os.environ["LINKDING_URL"].rstrip("/")
HEADERS = {"Authorization": "Token " + os.environ["LINKDING_TOKEN"],
"Content-Type": "application/json"}
def call(path, method="GET", body=None):
data = json.dumps(body).encode() if body else None
req = urllib.request.Request(BASE + path, data, HEADERS, method=method)
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)
for row in csv.DictReader(open(sys.argv[1], newline="", encoding="utf-8")):
if row["status_class"] != "gone": # errors can be a bad day, not a dead site
continue
check = call("/api/bookmarks/check/?url=" + urllib.parse.quote(row["url"], safe=""))
mark = check.get("bookmark")
if not mark:
continue
tags = set(mark["tag_names"]) | {"dead-link"}
if row["wayback_url"].startswith("http"):
tags.add("has-wayback")
call(f"/api/bookmarks/{mark['id']}/", "PATCH", {"tag_names": sorted(tags)})
print("tagged", row["url"])
Terminal window
LINKDING_URL=https://linkding.example.lan LINKDING_TOKEN=your-token python3 tag_dead_links.py report.csv

Only gone rows get tagged. A timeout or a 5xx can be a bad afternoon on the other end, so those stay untouched until you re-run the check a week later. The script adds tags to what is already there and never removes a thing. Use the API for this pass, not for the bulk import: the importer handles batches, and one POST per link is slower.

Verify Before You Close the Old Tabs

  1. Compare counts. Linkding’s bookmark total should equal the “written” number from the summary, plus anything you already had in linkding before the import.
  2. Sort by oldest and spot-check the first ten. Their dates should match your oldest exports, not the import day.
  3. Search #source:pocket, #source:chrome and the rest. Each source tag should return a plausible count.
  4. Search #dead-link and open three results to confirm they’re dead.
  5. Search one of your folder tags, like #home-lab, and confirm the folder’s contents showed up.
  6. Check the unread filter. Your Pocket “unread” and Pinboard “toread” items should be there.

If the count is off by a handful, check the import output for failed lines. Those are usually duplicates that differ only in ways linkding’s key does not see.

Keep the original exports somewhere safe for a month. Your 2 AM self will thank you.

The SumGuy Take

The normalizer is about 200 lines because the problem is small. The hard part is not parsing but matching the target’s rules: same URL key, same tag limits, same flag semantics. Get those right and the importer does the dedup for you. Skip the deleting. Disk is cheap, and a link tagged dead-link with a Wayback copy beats a link you can’t remember saving.

Common Questions

Does linkding import Chrome bookmark folders as tags?

No. Linkding ignores the folder structure in a Netscape HTML file and has no folders of its own. Chrome exports carry no tags, so every folder name is lost on import. Convert folders to tags before importing, which is what normalize.py --folder-tags does for you.

Can I still export my Pocket bookmarks in 2026?

No. Mozilla shut Pocket down on July 8, 2025, and the export page closed in late 2025. If you saved a CSV export before then, linkding can’t read it directly, but normalize.py converts it to Netscape HTML. If you never exported, the data is unrecoverable.

Will importing the same file twice into linkding create duplicates?

No. Linkding matches on a normalized URL (lowercase scheme and host, trailing slashes stripped, query parameters sorted) and updates the existing bookmark with the file’s title, description, dates, and flags. A repeat import is safe, and a second copy of the same URL inside one file is skipped and counted as failed.

How do I keep dead bookmarks without cluttering linkding?

Tag them dead-link and leave them in place. Dead entries only show up when you search or click the tag, so they stay out of the way until you want them. A second tag, has-wayback, marks the ones with an archived copy, so you can find those first.

Does Karakeep import tags from a Netscape HTML file?

Yes. Karakeep preserves titles, tags, and add dates from a Netscape HTML import and puts every imported bookmark into an automatically created list. It cannot import a subset, so trim dead links from the file first. Expect AI tagging and crawling to run in the background afterward.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Next Post
OpenCost on k3s: What Your Pods Cost

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts