100 million BullMQ jobs a day, and the dashboard we needed to watch them
100 million BullMQ jobs a day, and the dashboard we needed to watch them
Monest's Scale and Initial Approach
At Monest, everything happens after the webhook on BullMQ. Messages arrive, get queued, an agent drafts a reply, and it gets sent. Today that is more than 100 million jobs a day. From the start, the team recognized they needed visibility into these jobs - leading to the adoption of bull-board.
bull-board: Good Until the Team Grew
We installed bull-board. It was nice. It worked. It did exactly what it says. However, on day one the first problem appeared: it had no login. This was easy to fix - we put it behind the VPN and added a user and a password in front of it. Later, as the team grew, new questions emerged:
- What if someone deletes a job? How do we know who?
- Who paused this queue, and why?
- Can developers look, while only tech leads retry or remove?
None of that was possible with a single shared password - everyone became everyone.
Moving to Taskforce (BullMQ Dashboard)
We switched to Taskforce, the dashboard from the BullMQ team. The alerts were genuinely useful. But day-to-day, the UI felt slow and clunky to the people who lived in it. So I built the one I wanted: Bullpane.
What Bullpane Does
- Priority queues - The queues that need attention are at the top. When you open it, you see which queues are breaking rules (e.g.,
25% failed · 15m > 3%,780 waiting > 200) before anything else. - Quick navigation - Press
โKto jump to any queue on any connection. - Deep search - Search inside job data. For example, "Which job had order 81723?" Unlike bull-board, search here matches the payload (bounded and resumable), so it never blocks Redis.
- Group visualization - BullMQ Pro groups are displayed as distinct categories: waiting, limited, maxed, paused, each showing concurrency and rate limits per group.
- Team-specific roles - Developers get viewer, tech leads get operator, and the server enforces these permissions on every API call, not just by hiding buttons.
- Audit log - Track who paused the queue, who removed the job, who hit retry-all, and from which IP.
- Export & collaboration - Append-only, CSV export available. SSO via OIDC or SAML eliminates manual invite chains.
- Flexible alerting - Alerts can be routed per queue, per folder, or per whole connection to Slack or a webhook. Folders let teams group queues the way they think about them.
Performance Optimization for Redis
At 100M jobs a day, a dashboard that slows down the workload it watches is worse than no dashboard. The challenge: how to monitor without hammering Redis.
Key Design Principles
- No KEYS, ever. Discovery uses a bounded
SCAN, cached for efficiency. - One round trip per read: Counts and pagination are handled by Lua scripts over
EVALSHA. This keeps latency low even under heavy load. - Payload truncation inside Redis. A page of 200 jobs with 1 MB payloads costs ~7 ms; a search over 1,000 such jobs takes ~8 ms.
- Writes go through the official bullmq client. I never reimplemented its Lua logic.
Production Considerations
A read-only mode was implemented for early production days. The numbers and hostile-load tests are in the repository (STRESS-TEST.md). You can try it live at demo.bullpane.com (read-only, simulating a busy company). To run locally:
git clone https://github.com/madmorett/bullpane
cd bullpane
docker compose up -d
Then set your Redis URL and start with BULLPANE_READ_ONLY=true. The core is free and MIT-licensed, requiring no account or login - making it ideal for testing on existing queues.
If you try it on your queues, please reach out at h****@bullpane.com or file an issue on GitHub.
Comments
No comments yet. Start the discussion.