100 million BullMQ jobs a day, and the dashboard we needed to watch them
DEV Community

100 million BullMQ jobs a day, and the dashboard we needed to watch them

100 million BullMQ jobs a day, and the dashboard we needed to watch them

Monest's Scale and Initial Approach

At Monest, everything happens after the webhook on BullMQ. Messages arrive, get queued, an agent drafts a reply, and it gets sent. Today that is more than 100 million jobs a day. From the start, the team recognized they needed visibility into these jobs - leading to the adoption of bull-board.

bull-board: Good Until the Team Grew

We installed bull-board. It was nice. It worked. It did exactly what it says. However, on day one the first problem appeared: it had no login. This was easy to fix - we put it behind the VPN and added a user and a password in front of it. Later, as the team grew, new questions emerged:

  • What if someone deletes a job? How do we know who?
  • Who paused this queue, and why?
  • Can developers look, while only tech leads retry or remove?

None of that was possible with a single shared password - everyone became everyone.

Moving to Taskforce (BullMQ Dashboard)

We switched to Taskforce, the dashboard from the BullMQ team. The alerts were genuinely useful. But day-to-day, the UI felt slow and clunky to the people who lived in it. So I built the one I wanted: Bullpane.

What Bullpane Does

  • Priority queues - The queues that need attention are at the top. When you open it, you see which queues are breaking rules (e.g., 25% failed · 15m > 3%, 780 waiting > 200) before anything else.
  • Quick navigation - Press โŒ˜K to jump to any queue on any connection.
  • Deep search - Search inside job data. For example, "Which job had order 81723?" Unlike bull-board, search here matches the payload (bounded and resumable), so it never blocks Redis.
  • Group visualization - BullMQ Pro groups are displayed as distinct categories: waiting, limited, maxed, paused, each showing concurrency and rate limits per group.
  • Team-specific roles - Developers get viewer, tech leads get operator, and the server enforces these permissions on every API call, not just by hiding buttons.
  • Audit log - Track who paused the queue, who removed the job, who hit retry-all, and from which IP.
  • Export & collaboration - Append-only, CSV export available. SSO via OIDC or SAML eliminates manual invite chains.
  • Flexible alerting - Alerts can be routed per queue, per folder, or per whole connection to Slack or a webhook. Folders let teams group queues the way they think about them.

Performance Optimization for Redis

At 100M jobs a day, a dashboard that slows down the workload it watches is worse than no dashboard. The challenge: how to monitor without hammering Redis.

Key Design Principles

  • No KEYS, ever. Discovery uses a bounded SCAN, cached for efficiency.
  • One round trip per read: Counts and pagination are handled by Lua scripts over EVALSHA. This keeps latency low even under heavy load.
  • Payload truncation inside Redis. A page of 200 jobs with 1 MB payloads costs ~7 ms; a search over 1,000 such jobs takes ~8 ms.
  • Writes go through the official bullmq client. I never reimplemented its Lua logic.

Production Considerations

A read-only mode was implemented for early production days. The numbers and hostile-load tests are in the repository (STRESS-TEST.md). You can try it live at demo.bullpane.com (read-only, simulating a busy company). To run locally:

git clone https://github.com/madmorett/bullpane
cd bullpane
docker compose up -d

Then set your Redis URL and start with BULLPANE_READ_ONLY=true. The core is free and MIT-licensed, requiring no account or login - making it ideal for testing on existing queues.

If you try it on your queues, please reach out at h****@bullpane.com or file an issue on GitHub.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.