Importing from Browsertrix
If your WACZs live in a Browsertrix account (Webrecorder’s hosted crawler), indice import browsertrix downloads them into <home>/archive and indexes them as durable local sources.
Credentials come from the environment, never the command line, so they don’t show up in the process list:
export BROWSERTRIX_USER='you@example.org'read -rs BROWSERTRIX_PASSWORD; export BROWSERTRIX_PASSWORD # prompts, no echo# or, instead of user/password: export BROWSERTRIX_TOKEN='<a JWT>'Then import — preview first with --dry-run, then pull for real:
# everything you've QA'd in your (only) org, into ~/webarchiveindice import browsertrix --home ~/webarchive --dry-runindice import browsertrix --home ~/webarchive
# just one collection (by id, slug, or name) → a matching indice collectionindice import browsertrix --collection us-govarchive --home ~/webarchive
# a single crawlindice import browsertrix --crawl <item-id> --home ~/webarchive- QA’d crawls only, by default. Browsertrix lets a reviewer rate a crawl (
reviewStatus); indice imports only reviewed crawls so you publish vetted content. Add--include-unreviewedto import everything, or--min-review <1-5>for a rating threshold. A single named--crawlis always imported. When crawls are skipped for this reason, indice says so. - Selection.
--collection <ID|SLUG|NAME>limits to one Browsertrix collection;--crawl <ID>to a single archived item; neither imports the whole org.--org <SLUG>picks the org when your account has more than one. - Incremental. Re-running skips crawls already imported (matched by content hash), so syncing an account is cheap;
--forcere-imports anyway. - Durable by default. WACZs are downloaded into
<home>/archive/<item-id>/(a subfolder per Browsertrix item, so items can’t clash on a shared filename), because Browsertrix’s presigned URLs expire after ~48h — a downloaded copy keeps replay working long-term.--host <URL>targets a self-hosted Browsertrix (default ishttps://app.browsertrix.com). --stream(index-only footprint). Instead of downloading, index the WACZ in place from Browsertrix and store only its stable identity, not a copy. Since presigned URLs expire, indice re-resolves a fresh one on demand — soserveneeds the sameBROWSERTRIX_*credentials to replay these crawls (they show a 503 otherwise). Good for a self-hoster who wants search over their own crawls without keeping the bytes; download (the default) is better for a durable, offline, or shared library.- Grouping. Importing a
--collectiongroups its crawls into an indice collection of the same name.--into <NAME>overrides that name (and is the way to group an org-wide or single---crawlimport, which otherwise land as individual collections).--limit <N>caps how many are imported;--dry-runlists them without downloading.
Both importers are also available in management mode as browse-and-import wizards on the accession desk, using the server’s configured credentials.