- mention @ArchiveBox in any thread β saves all URLs found in the last 10 messages
- DM @ArchiveBox bot with any text β archives every URL in your DM
- Add @ArchiveBox to group chats β optionally archive all URLs shared in groups it's added to
- Auto-tags URLs with source info β
slack,bob,reading-list (connector, user, channel name)
π³ Docker β ArchiveBox URL + API key β chat provider
git clone https://github.com/ArchiveBox/archivebox-chat-bot.git
cd archivebox-chat-bot
docker compose pull
docker compose up -d
- Open Chatbot Admin Console β choose your password β ArchiveBox server URL + API key.
- Choose a provider below β connect ArchiveBox Bot.
- Pick conversations β New URLs, Saved URLs, groups, permissions.
- Data:
./archivebox-chat-bot/ Β· docker-compose.yml.
- Optional password preset: uncomment
ADMIN_PASSWORD to skip the first-run screen.
- Updates: run
docker compose pull && docker compose up -d again.
- Images:
archivebox/archivebox-chat-bot:latest on Docker Hub, or ghcr.io/archivebox/archivebox-chat-bot:latest, for amd64 and arm64. Each successful main build publishes latest, build-<CI run number>, and sha-<full commit SHA> after tests and container startup checks pass. No manual version bump is needed; Check β Run workflow also rebuilds main.
- Local development: uncomment
build: . in Compose and run docker compose up -d --build.


- Reachable links: Tailscale for personal use; a domain + HTTPS for internet access.
- Localhost / 127.0.0.1: allowed with a warning; set Public URL so links work on other devices.

- Connect: choose Create Slack app in the Chatbot Admin Console β install the preconfigured app β follow the two screenshot guides to paste your Bot token (
xoxb-β¦) and App token (xapp-β¦) β Save & connect.
- Use: invite the bot to channels; mention it in a thread or send a DM.
- Destinations: Save & connect creates New URLs + Saved URLs automatically. Socket Mode needs no public webhook.
| 1 Β· Set up the Slack app | 2 Β· Configure the bot | 3 Β· Mention in a thread |
 |
 |
 |
| Or send a DM | Browse Saved URLs |
 |
 |
Where to get Slack tokens
| Create the preconfigured app | Allow it in your workspace |
 |
 |


- Connect: create an app β Bot β Reset Token β paste token into the Chatbot Admin Console.
- Enable: Bot β Message Content Intent β Save Changes β Find my bot β Add to Discord β select your server β Save & connect.
- Use: DM links, mention the bot in channels/threads, or use
/archivebox save, search, auto, status, help.
- Destinations: New URLs + Saved URLs are created automatically; choose existing channels under bot preferences. No public webhook.
| 1 Β· Get your bot token | 2 Β· Enable Message Content | 3 Β· Add to your server |
 |
 |
 |
Create the app Β· install both bots

Official setup Β· Message Content Intent
| 4 Β· Choose channels & preferences | 5 Β· Browse Saved URLs |
 |
 |
Live Cabbage check (2026-10-08): A channel mention, capture-bot DM, and AI DM were handled. All tested Ars Technica URLs resolve to snapshots whose captured response is 403 Forbidden; replay shows that error instead of article content.
| Capture-bot channel mention | Capture-bot DM | Saved URL cards | AI DM reply |
 |
 |
 |
 |

| 1 Β· Copy bot email | 2 Β· Copy API key | 3 Β· Browse Saved URLs |
 |
 |
 |
Live Cabbage check (2026-10-08): The stream mention and AI DM flows were handled, but both Ars Technica URLs already had snapshots showing 403 Forbidden. The AI reply reported that ONLY_NEW skipped a retry, so these results do not show captured article content.
| Capture-bot stream mention | Saved URL results | AI DM reply |
 |
 |
 |

- Connect: BotFather β /newbot β copy token β add to group.
- Groups: disable bot privacy, then remove/re-add the bot for group-wide capture.
- Commands:
/save, /search, /auto, /status, /help; polling needs no public webhook.
| 1 Β· Add to your group | 2 Β· Choose bot preferences | 3 Β· Save references |
 |
 |
 |
BotFather Β· names & group privacy




Live Cabbage check (2026-10-08): A real WhatsApp Agent DM used /archivebox save with two Ars Technica URLs. The DM confirmed the request, and the Cabbage console recorded two saved URLs and a completed message. Both snapshots contain Ars Technica's HTTP 403 Forbidden response, not article content.
| Real WhatsApp DM | Cabbage saved result |
 |
 |

| 1 Β· Configure the bot | 2 Β· Browse Saved URLs |
 |
 |
Live Cabbage check (2026-10-08): In the signed-in Lounge UI, ArchiveBoxTester mentioned ArchiveBoxCapture in #new-urls and sent two Ars Technica URLs in a capture-bot DM. Both captures completed, and the saved channel shows three result cards: April channel snapshot 06ac7719f3997373800022139b4a76f3, plus the April and October DM snapshots 06ac76f04b0473098000a102ed4bd265 and 06ac76f04b047f6780006432632c6c0f. Each result is titled 403 Forbidden; the snapshots contain the origin's error response, not article content. The IRC AI bot was configured through the console at ergo:6667 (plain IRC, no TLS), after its old 127.0.0.1:16667 endpoint failed. Its request created snapshots 06ac76358bae7ed78000c4cfe72049be and 06ac76358baf7cb88000cdb323aa1f0e, also HTTP 403. The AI used the normal CLI's --no-only-new option because ONLY_NEW skipped URLs with existing snapshots. In the capture job records, num_snapshots remained 0 even though their results listed the snapshot IDs.
| Channel mention | Capture-bot DM | Saved snapshot result |
 |
 |
 |

| 1 Β· Enable Full Disk Access | 2 Β· Allow Messages Automation |
 |
 |
Docker β Mac prerequisites
- SSH key + verified
known_hosts, readable by container UID 1000.
- Set the Mac's absolute
imsg path; Compose does not provision the Mac or SSH access.

| 1 Β· Create a Page | 2 Β· Generate token & configure webhook |
 |
 |
Meta webhook
- Register
https://<chatbot-host>/connections/<connection-id>/capture/webhook with your matching verify token.
- Subscribe the Page to message events; use
/ai/webhook for ArchiveBox AI Bot.
Live Cabbage check (2026-10-08): A real Facebook Messenger Page conversation sent two Ars Technica URLs to the connected ArchiveBox Bot Page. The webhook returned HTTP 200, and Cabbage recorded two saved URLs and a completed capture. Ars Technica returned HTTP 403 Forbidden for both snapshots, so they contain the error response rather than article content.
| Real Messenger conversation | Cabbage saved result |
 |
 |

Live Signal check (2026-10-08): Real Signal Desktop messages in Note to Self and the dedicated ArchiveBox Cabbage Test group produced eight saved URLs through Beeper. The DM and group capture jobs finished, and their own snapshots were verified as sealed. Ars Technica returned 403 Forbidden: the saved outputs contain that response, not the article bodies. The group screenshot below shows the saved notification for its own capture, matched to its snapshot ID.
| Signal DM and queue acknowledgement | Signal group submission |
 |
 |


Docker β Beeper address
- Use the Beeper server's reachable network address; container
localhost points at the bot itself. Beeper runs separately; iMessage needs a Mac.
- Connect: dedicated mailbox β IMAP server + address + app password β Save & connect. AgentMail can create an unclaimed, receive-only inbox; use
imap.agentmail.to, port 993, TLS, the inbox address, and its API key as the password.
- Send / forward / CC β save links from the subject, body, quoted conversation, and text/HTML/
.eml attachments.
- Tags:
email + sender's name Β· Results: ArchiveBox + Activity.
- Inbox: checked every 30 seconds; new mail only by default. Optional first-connection import of existing mail.
- Inbound only: no SMTP or replies; preserves read/unread flags. PDF, Office, and image attachments are not inspected or uploaded.
Live Gmail β AgentMail check (2026-10-08): Two real Gmail messages with two Ars Technica URLs each were read from AgentMail over IMAP, parsed, and saved by Cabbage. The UID 3 and UID 4 jobs completed with four sealed ArchiveBox snapshots and non-empty saved outputs, including WACZ, screenshot, and hash-manifest files. Ars Technica returned 403 Forbidden for these requests, so the captures do not contain the article bodies.



Google app passwords Β· Gmail IMAP Β· iCloud IMAP Β· Fastmail IMAP
- Invite the bot β send a message β select the group in the Chatbot Admin Console.
- Enable Archive every link per group or across joined groups; choose administrators under Permissions.
| Command |
Action |
/archivebox save <URLs> |
Capture links |
/archivebox search <words> |
Find saved pages |
/archivebox auto on / off |
Toggle this group's automatic capture |
/archivebox status |
Check ArchiveBox |
/archivebox help |
Show commands |
Capture & provider details
- Capture depth 0, selected persona.
- History: provider API where available, otherwise messages received while connected.
- Zulip: commands in DMs or after a mention. IRC/iMessage: commands as ordinary message text.
- Reactions/media depend on the provider; IRC/iMessage lack reliably targeted reactions, Messenger Pages use text replies.
Deployment & updates
- Include ArchiveBox: uncomment the
archivebox service in docker-compose.yml, then run docker compose up -d --build.
- ArchiveBox inside Compose:
http://archivebox:5797; for a separately hosted server, use its reachable URL.
- Public URL is the address readers open; split API/admin hosts are discovered automatically.
- Back up
./archivebox-chat-bot/; run one bot service per data directory.
- Update:
git pull --ff-only && docker compose up -d --build.
Credentials & recovery
- Activity β jobs, sessions, Recover answer, Chatbot Admin Console password.
- Queued jobs retain their original server/connection; review uncertain delivery before retrying.
- Blank credential fields preserve secrets. Protect the data directory; use HTTPS for remote access.
- Revoke credentials to remove access. Removing bot data preserves snapshots and chat messages.
Development
uv sync
cd connectors && npm ci && npm run build && cd ..
uv run archivebox-chat-bot
uv run python -m pytest -xq tests/local
uv run ruff check .
- Python: shared capture/agent engine, queue, permissions, Chatbot Admin Console.
- Transports: Beeper API, Chat SDK, Slack/Zulip APIs, discord.py, pydle, imsg.
- Live tests:
tests/test_*_live.py. Chat and setup screenshots are our sessions. Provider logos are official website assets; IRC uses the Libera.Chat network logo, and the bundled Messenger logo comes from Messenger.
Applies to ArchiveBox Bot and ArchiveBox AI Bot. Updated 2026-10-08.
- Self-hosted: each deployment is run by its operator, who controls access, configuration, storage, and availability. Installing a Discord app does not provision an ArchiveBox server.
- Authorized use: archive only content you are entitled to access and preserve; respect applicable laws, content rights, and your chat provider's rules. Operators must inform participants when automatic capture is enabled.
- Your responsibility: protect credentials and backups, choose who can use the bots, and review AI actions and results. Captures and AI answers may be incomplete or incorrect.
- License: the software is provided under the MIT License, including its warranty disclaimer and limitation of liability. No hosted service, uptime guarantee, or paid subscription is included.
- Stop using it: disconnect the bots and revoke their credentials. Existing archives, sessions, messages, and backups require separate deletion by their operators.
- Help: project issues Β· ArchiveBox community. For a particular deployment or data-removal request, contact the person or organization operating that bot.
Applies to ArchiveBox Bot and ArchiveBox AI Bot. Updated 2026-10-08.
- Data received: message text, sender names/identifiers, conversation names/identifiers, and message/thread identifiers made available by the connected provider. Email capture also reads message subjects and supported attachment text.
- Local storage: the operator's data directory stores credentials, configuration, recent conversation history, and job requests/results/errors. Received messages can enter history even when capture permissions prevent archiving. Active history retains up to 200 messages per connection, bot, and conversation; jobs, disconnected-account data, and backups have no automatic expiry.
- Capture: extracted URLs and source tags (provider, sender, channel/group) go to the configured ArchiveBox server. That server contacts archived websites and any services enabled by its operator. Capture cards, screenshots, favicons, search results, and replies can be posted to the configured chat destinations and become visible to their members.
- Optional AI: trusted users' requests and conversation context go to the configured ArchiveBox/OpenCode service and its configured model providers and tools. These services have their own processing and retention policies. AI is disabled until configured and enabled.
- Security: the database uses restrictive filesystem permissions; stored credentials are not encrypted by the application. The Chatbot Admin Console password is hashed. Operational logs can contain errors and exceptions. Operators must protect the host, data directory, backups, and remote connections.
- No built-in advertising or analytics: the bot does not send usage analytics to ArchiveBox maintainers. Chat providers, hosting services, archived websites, and optional AI services process data under their own policies.
- Access and deletion: contact the deployment operator to request access, correction, or deletion of retained data and backups. Disconnecting/revoking credentials stops further access; deleting bot data does not delete ArchiveBox snapshots, OpenCode sessions, or messages already posted to chat. Those must be removed separately in their respective systems.
- Questions: contact the deployment operator first. Project support is available through ArchiveBox community or issues; do not post private messages or credentials in public support channels. Revisions to this policy appear here with an updated date.
- DM / @mention β capture, tag, organize, and verify.
- Follow up β include recent messages and previous replies.
- ArchiveBox β Agent β inspect each new session and tool results.
- Existing OpenCode β providers, credentials, tools, session database; tasks survive restarts.
- ArchiveBox AI Bot in the Chatbot Admin Console β connect a second bot/account.
- Trusted people β choose who can run tasks.
- Agent preferences β optional prompt β enable.

- Connect: choose Create Slack app for ArchiveBox AI Bot β install the second preconfigured app β paste its Bot token and App token β Save & connect. Native agent status and Stop controls appear in Slack.
- Shown: plan β approval β capture HTTP caching references β tag and verify.
| 1 Β· Configure ArchiveBox AI Bot | 2 Β· Complete an archive task |
 |
 |

- Connect: create a second Discord app named ArchiveBox AI Bot β repeat the Discord setup β choose Trusted people.
- DM / mention: capture, tag, and verify with your existing ArchiveBox agent; follow up in the same DM or thread.
- ArchiveBox β Agent: open the session and its tool results.
Choose who can use ArchiveBox AI Bot



- Connect: second BotFather token; add ArchiveBox AI Bot to the group.
- Shown: a real Telegram DM asks the Cabbage server for its collection path and snapshot count; OpenCode runs the ArchiveBox command and replies in Telegram.
| 1 Β· Connect the AI bot | 2 Β· Query the live collection |
 |
 |

- Connect: second nickname/account on the same server.
- Shown: find a missing asyncio reference β capture it β answer from saved HTML.
| 1 Β· Connect the AI bot | 2 Β· Search archived sources |
 |
 |
| IRC AI connection after host correction | Real Ars request | AI result |
 |
 |
 |

- Connect: separate Agent API key or linked account; choose trusted people.
- Shown: inspect existing archives β identify a missing Fetch reference β propose capture/tagging.
| 1 Β· Connect the AI bot | 2 Β· Review the proposed plan |
 |
 |