Skip to content
Go back

Crawl4AI: Feed Your RAG From the Web

By KingPin 13 min read
Crawl4AI: Feed Your RAG From the Web
Contents

Your RAG Is Only as Good as What You Feed It

Crawl4AI is the right tool when the docs you want live behind JavaScript, span hundreds of pages, or come wrapped in so much navigation chrome that your chunks are mostly sidebar. It is the wrong tool when the site has a sitemap and plain static HTML, because requests plus trafilatura does that job in 20 lines and starts instantly.

This post builds the full path a home-labber would run: Crawl4AI crawls a docs site into clean markdown, a script chunks it by heading, Ollama embeds it with nomic-embed-text, Qdrant stores it, and a query script pulls back the closest passages. Everything below was run end to end against Crawl4AI 0.9.4 (released September 2026, Apache-2.0 licensed, as of October 2026) with Qdrant 1.19.1.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/llm/crawl4ai-rag-pipeline

Using a headless browser to read a static page is like hiring a forklift to move a couch. Technically it works, but your neighbors will have questions. So we start with the exit ramp.

Do You Even Need Crawl4AI?

Check these first. Each one is cheaper than a browser fleet.

  1. The site publishes llms.txt or a markdown export. Many docs sites do now. Fetch https://docs.example.com/llms.txt and you have a curated index of pages, often with .md links. No crawler needed.
  2. The site has an API. GitHub, Discourse, Read the Docs, and most wikis expose content as JSON or raw markdown. Use that.
  3. Static HTML plus a sitemap. Pull sitemap.xml, fetch each URL with httpx, run trafilatura.extract(). Done.
  4. The docs are a git repo. Clone it. The source markdown is cleaner than anything a crawler produces.

Reach for Crawl4AI when you hit one of these: client-side rendered pages (React or Vue docs that return an empty shell without JavaScript), sites with no sitemap where you must follow links, or pages where you want the page chrome stripped by a relevance filter instead of hand-written selectors.

The Stack: Three Containers and a Venv

Qdrant stores vectors. Ollama runs the embedding model. Crawl4AI runs from a Python venv (the Docker REST server comes later).

compose.yaml
services:
qdrant:
image: qdrant/qdrant:v1.19.1
ports:
- "6333:6333"
volumes:
- qdrant_data:/qdrant/storage
restart: unless-stopped
ollama:
image: ollama/ollama:latest
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
restart: unless-stopped
volumes:
qdrant_data:
ollama_data:

Bring it up and pull the embedding model:

Terminal window
docker compose up -d
docker compose exec ollama ollama pull nomic-embed-text

nomic-embed-text produces 768-dimension vectors. I checked: a call to /api/embed returns a list of 768 floats per input. That number matters in a minute, because Qdrant refuses vectors that do not match the collection size. Ollama also has embeddinggemma (300M) if you want something newer. Swap the name and the dimension and nothing else changes.

Now the crawler:

Terminal window
python3 -m venv .venv && . .venv/bin/activate
pip install crawl4ai qdrant-client httpx
crawl4ai-setup # downloads the Playwright browser

Crawl a Docs Site

Crawl4AI’s deep crawl runs a BFS (breadth-first) strategy over links, with filters that decide which URLs are worth visiting. Here is the whole script:

crawl.py
import asyncio
import hashlib
import json
import sys
from pathlib import Path
from crawl4ai import (
AsyncWebCrawler,
BFSDeepCrawlStrategy,
BrowserConfig,
CacheMode,
CrawlerRunConfig,
DefaultMarkdownGenerator,
PruningContentFilterLXML,
)
from crawl4ai.deep_crawling import DomainFilter, FilterChain, URLPatternFilter
START_URL = sys.argv[1] if len(sys.argv) > 1 else "https://docs.example.com/"
MAX_PAGES = int(sys.argv[2]) if len(sys.argv) > 2 else 50
OUT = Path("pages")
async def main() -> None:
host = START_URL.split("/")[2]
config = CrawlerRunConfig(
deep_crawl_strategy=BFSDeepCrawlStrategy(
max_depth=2,
max_pages=MAX_PAGES,
filter_chain=FilterChain(
[
DomainFilter(allowed_domains=[host]),
URLPatternFilter(patterns=["*/changelog/*", "*/blog/*"], reverse=True),
]
),
),
markdown_generator=DefaultMarkdownGenerator(
content_filter=PruningContentFilterLXML(threshold=0.48, threshold_type="fixed")
),
cache_mode=CacheMode.ENABLED,
check_robots_txt=True,
semaphore_count=2,
mean_delay=1.0,
max_range=1.0,
verbose=False,
user_agent="MyHomelabRAG/1.0 (+https://example.com/bot)",
)
OUT.mkdir(exist_ok=True)
async with AsyncWebCrawler(config=BrowserConfig(headless=True, verbose=False)) as crawler:
results = await crawler.arun(START_URL, config=config)
kept = 0
for r in results:
if not r.success:
print(f"skip {r.url}: {r.error_message}")
continue
md = r.markdown.fit_markdown or r.markdown.raw_markdown
if len(md) < 200:
continue
name = hashlib.sha1(r.url.encode()).hexdigest()[:12]
(OUT / f"{name}.md").write_text(md, encoding="utf-8")
(OUT / f"{name}.json").write_text(json.dumps({"url": r.url}))
kept += 1
print(f"crawled {len(results)} pages, kept {kept}")
asyncio.run(main())

Run it with python crawl.py https://docs.example.com/ 50. A few things in there earn their keep.

max_pages and max_depth are your seatbelt. A docs site with a version switcher can have thousands of URLs. Without a cap, your 2 AM self gets to learn what a 40,000-page crawl does to a Pi.

DomainFilter keeps the crawl on the host. include_external already defaults to off. The filter adds an explicit allow-list: it lets through the host and its subdomains, and blocks sibling domains such as the bare parent domain. The URLPatternFilter with reverse=True excludes matching URLs, which is how you skip changelogs and blog posts that bloat the index with stale version noise.

The result of arun with a deep crawl strategy is a list of results (when stream=False, the default). Each one carries .markdown, an object with raw_markdown, fit_markdown, markdown_with_citations, and references_markdown. It also prints as the raw markdown if you treat it as a string.

Raw Markdown vs Fit Markdown

raw_markdown is the whole page converted to markdown. That includes nav menus, footers, “Edit this page” links, and cookie banners. Embed that and your vector index fills with chunks that all say “Home / Docs / API / Pricing”.

fit_markdown is the same page after a content filter removes low-value blocks. The filter lives on the markdown generator:

In 0.9.4 the older PruningContentFilter still works but prints a deprecation warning telling you to switch to the LXML variant, so the code above uses that. This is the kind of change you hit when you copy a 0.4-era tutorial.

On one docs page from a test crawl, the raw markdown was 26,132 characters and the fit version 18,543. That is 29% less text to embed, and the cut text was menus and link lists. Pruning is heuristic, so it sometimes drops a short but useful block. Spot-check a few pages before trusting it with 5,000.

Be Polite or Get Blocked

Crawlers get banned for being rude, and rude is the default here. Check the knobs against what you want:

For large batches with arun_many, Crawl4AI also ships a RateLimiter (with exponential backoff on 429 and 503) and a MemoryAdaptiveDispatcher that slows down when RAM gets tight. You pass the dispatcher as arun_many(urls, config=..., dispatcher=...). Deep crawls like the one above use the per-config settings instead.

Robots.txt is a politeness convention, not a legal shield. Terms of service and copyright still apply to what you do with the content. Crawling a project’s own public docs for a private index is low risk, though terms of service vary. Re-publishing someone else’s content is a different story.

JavaScript-Rendered Pages

This is where the browser earns its RAM. For a page that renders content client-side, tell Crawl4AI what “loaded” looks like:

js_page.py
config = CrawlerRunConfig(
wait_for="css:.docs-content", # block until this selector exists
scan_full_page=True, # scroll to trigger lazy loading
delay_before_return_html=0.5,
)

wait_for accepts a css: selector or a js: expression that returns true. The default wait_until is domcontentloaded, which returns before most single-page apps finish drawing. If your markdown comes back as 40 characters, the page was not ready. Add a wait_for.

Caching: Stop Hammering the Site While You Debug

You will run this script 15 times while tuning chunk sizes. CrawlerRunConfig defaults to CacheMode.BYPASS, which always fetches fresh. The script above uses CacheMode.ENABLED, so the second run reads pages from the local cache and the site sees nothing. The other modes are DISABLED, READ_ONLY, and WRITE_ONLY.

Use ENABLED while iterating. Switch to BYPASS when you want a fresh crawl for real.

Chunk, Embed, Upsert

Crawl4AI ships chunking strategies, but for docs you want headings. A heading is a topic boundary the author already drew for you. Splitting on #, ##, and ### keeps each chunk about one thing, and the heading text goes along with it so the embedding knows the subject.

ingest.py
import json
import re
import uuid
from pathlib import Path
import httpx
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, PointStruct, VectorParams
OLLAMA = "http://localhost:11434"
MODEL = "nomic-embed-text" # 768 dimensions
COLLECTION = "docs"
MAX_CHARS = 1500
def chunk(md: str) -> list[tuple[str, str]]:
parts = re.split(r"(?m)^(?=#{1,3} )", md)
out = []
for part in parts:
part = part.strip()
if len(part) < 80:
continue
heading = part.splitlines()[0].lstrip("# ").strip()
buf = ""
for para in part.split("\n\n"):
if buf and len(buf) + len(para) > MAX_CHARS:
out.append((heading, buf))
buf = f"{heading}\n\n"
buf += para + "\n\n"
out.append((heading, buf.strip()))
return out
def embed(texts: list[str]) -> list[list[float]]:
r = httpx.post(f"{OLLAMA}/api/embed", json={"model": MODEL, "input": texts}, timeout=120)
r.raise_for_status()
return r.json()["embeddings"]
def main() -> None:
qd = QdrantClient(url="http://localhost:6333")
if not qd.collection_exists(COLLECTION):
qd.create_collection(COLLECTION, vectors_config=VectorParams(size=768, distance=Distance.COSINE))
total = 0
for md_file in sorted(Path("pages").glob("*.md")):
url = json.loads(md_file.with_suffix(".json").read_text())["url"]
chunks = chunk(md_file.read_text(encoding="utf-8"))
if not chunks:
continue
vectors = embed([f"search_document: {text}" for _, text in chunks])
points = [
PointStruct(
id=str(uuid.uuid5(uuid.NAMESPACE_URL, f"{url}#{i}")),
vector=vec,
payload={"url": url, "heading": heading, "text": text},
)
for i, ((heading, text), vec) in enumerate(zip(chunks, vectors))
]
qd.upsert(COLLECTION, points=points)
total += len(points)
print(f"{url}: {len(points)} chunks")
print(f"upserted {total} chunks")
if __name__ == "__main__":
main()

Three details to know:

On five docs pages from a test crawl, this produced 78 chunks, and a second run left the count unchanged.

Query It

query.py
import sys
import httpx
from qdrant_client import QdrantClient
question = " ".join(sys.argv[1:]) or "How do I limit how fast the crawler hits a site?"
r = httpx.post(
"http://localhost:11434/api/embed",
json={"model": "nomic-embed-text", "input": [f"search_query: {question}"]},
timeout=60,
)
r.raise_for_status()
qd = QdrantClient(url="http://localhost:6333")
hits = qd.query_points("docs", query=r.json()["embeddings"][0], limit=3).points
for h in hits:
print(f"{h.score:.3f} {h.payload['url']} [{h.payload['heading']}]")
print(" ", h.payload["text"][:200].replace("\n", " "), "\n")

query_points is the current search call. In qdrant-client 1.19.1 the old search method no longer exists, so tutorials that use it fail with an AttributeError.

With only five pages indexed, my results were mediocre. Questions about cache modes found the right quickstart section. A question about robots.txt returned a “How You Can Support” block, because the tiny corpus had nothing better. That is the honest result of a five-page index. Crawl the whole site before you judge retrieval quality. To answer questions, feed the top hits plus the question to a local LLM through the same Ollama server.

The REST API Route

If your ingest code is not Python, or you want the crawler on a different machine from the script, run the official Docker server. As of 0.9.4 the image is unclecode/crawl4ai:0.9.4, it listens on port 11235, wants --shm-size=1g (the README’s run command sets it, and Chromium is known to struggle on Docker’s small default), and requires an API token. Without a token it only answers requests from inside its own container.

Terminal window
export CRAWL4AI_API_TOKEN="$(openssl rand -hex 32)"
docker run -d -p 11235:11235 --name crawl4ai --shm-size=1g \
-e CRAWL4AI_API_TOKEN="$CRAWL4AI_API_TOKEN" \
unclecode/crawl4ai:0.9.4
curl -s http://localhost:11235/md \
-H "Authorization: Bearer $CRAWL4AI_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url": "https://docs.example.com/", "f": "fit"}'

/md returns one page as markdown, and f picks the filter (fit is the pruning one). For full control, POST /crawl takes urls, a browser_config, and a crawler_config, each shaped as {"type": "CrawlerRunConfig", "params": {...}}. I confirmed check_robots_txt and cache_mode pass through that way. The response holds a results list, and each result has the same markdown fields as the Python object.

Pin the tag so a pull next month does not change your API under you.

The SumGuy Take

Run the library for a one-time crawl of a site you care about, and put the REST container on the box if several scripts will share a browser. Turn on check_robots_txt, cap max_pages, cache while you iterate, and use fit markdown. Chunk by heading, keep the URL in the payload so answers can cite their source, and pin your embedding model, because switching models means re-embedding everything.

If the docs are static, skip all of it and use trafilatura. A browser you do not need is a cron job waiting to fail at 3 AM.

Common Questions

Does Crawl4AI respect robots.txt?

Only if you ask. The check_robots_txt option on CrawlerRunConfig defaults to False. Set it to True and Crawl4AI fetches and caches the rules, then returns a 403 result with “Access denied by robots.txt” for disallowed URLs instead of crawling them.

Can Crawl4AI run without a GPU?

Yes. Crawl4AI drives a headless Chromium browser and does not need a GPU. The Docker server wants 1 GB of shared memory (--shm-size=1g), and each open page costs RAM, so keep concurrency low on a small box. A GPU only matters for the embedding or LLM step, and nomic-embed-text runs fine on CPU.

Do I need an LLM to use Crawl4AI?

No. Crawling, markdown conversion, and the pruning and BM25 filters all run without any model. An LLM is optional, for features like LLMExtractionStrategy that pull structured fields out of a page. A RAG ingest pipeline needs only the crawler plus an embedding model.

Which embedding model should I use with Ollama?

Start with nomic-embed-text for a Crawl4AI and Ollama pipeline. The nomic-embed-text model outputs 768-dimension vectors, runs fine on CPU, and is the model this pipeline was tested with. embeddinggemma (300M parameters) is a newer option. Whatever you pick, set the Qdrant collection size to match the model’s dimension, and re-embed everything if you switch.

Often yes for personal use, but this is not legal advice. Crawling a project’s public docs for a private, personal index is low risk. Ignoring robots.txt, hammering servers, or re-publishing copyrighted content is where trouble starts. Read the site’s terms and honor its robots.txt file.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Previous Post
Grafana Dashboard Sprawl: 50 Panels to 10
Next Post
nginx vs Caddy: Measured on 2 Cores

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts