NAME
crawlnet — crawls the crypto web and pretrains a crypto-native LLM on it, from scratch.
SYNOPSIS
crawlnet [fees] → crawl → tokenize → pretrain → queen vn+1
DESCRIPTION
- crawl
- Crawlers drive real headless Chromium through the crypto web in chapters: whitepapers, docs, solana, forums, blogs, github. Every screen on this site is a real page loading right now.
- ingest
- Every page is canonicalized, deduplicated, relevance-gated and tokenized. Her tokenizer is trained on the crawl too.
- map
- The dataset is indexed by project and chapter. Chapters are finite target lists, so coverage can hit 100%.
- train
- When the treasury funds a run, a new queen is pretrained from random init on the dataset and nothing else, on rented GPUs. Each version is bigger and has read more. Weights ship public.
- spawn
- Burn $CRAWL to spawn your own crawler. Its pages are credited to your wallet, and every 12-hour round it works it earns a share of the owners' pool, up to its earn cap.
- capped
- A crawler earns up to 2x what it cost, then it stops earning (it keeps crawling). Spawn a new one to keep earning: more crawlers, more data for the queen.
- order
- Burn 10,000 $CRAWL and the crawlers read up to 50 pages of a site you name, within 3 hours: one host, robots.txt obeyed, never past a bot check. Automatic checks come first (robots.txt, a real page, no bot check, no scam, adult or gambling content); if the site passes, burn right away. Every order gets a public report of what the queen kept.
ECONOMICS
pump.fun creator fees land in the queen's wallet. 60% pays for crawl compute (browsers + models) and pretraining runs (GPU hours). 40% goes into a pool for crawler owners that closes every 12 hours, at 00:00 and 12:00 UTC: it is split equally per crawler that brought 25+ new pages in that round and paid in BNB, twice a day. Every BNB is on the ledger.
EXAMPLES
# spawn your own crawler, start it on your docs $ Spider Crawlers spawn --name merkle-weaver --seed https://docs.solana.com # what the crawlers have read so far $ curl -s $CRAWL_API/v1/stats # ask the queen $ curl -s -X POST $CRAWL_API/v1/queen/ask -d '{"question":"what is a validator?"}'
FILES
- /v1/live
- websocket: crawler state, frames of visible screens, pages, ledger
- /v1/web
- the dataset as a graph of pages and links
- /v1/treasury/ledger
- every BNB in and out
- /v1/orders/:id
- an ordered crawl: its status and every page read