What Running Kafka on VMs Taught Us About Systems Thinking
Celestina Amadi leads the Cloud Engineering team behind the infrastructure that powers Moniepoint's payment and savings products, building systems that move transaction and savings data reliably for millions of customers every day. She's also a Grafana Champion, a HashiCorp Ambassador, and an IBM Champion 2026. The article below takes us on a journey of how her team switched from running Kafka on VMs to Strimzi. The coolest part about this switch is that the VMs were working, yet she made the call to rebuild anyway because she could see where things were headed before they became a problem. โโโโโโโโโโ Why We Moved to Strimzi We moved to Strimzi because we looked honestly at how we were running Kafka and made a deliberate call: this does not scale, and we can do better. That kind of decision is actually harder than reacting to a crisis. Incidents are obvious - something breaks, and you fix it. Proactive architectural change requires you to see clearly enough to act on discomfort before it becomes a disaster. It requires you to say "this is working, but not well enough" and mean it. This is the story of what we saw, what we built, and what it taught us about thinking in systems. How We Were Running Kafka Before Our Kafka setup started the way most things do in a fast-moving engineering team: pragmatically. We needed CDC pipelines. Moving data from Database A to Kafka to Database B, from Database C to Kafka to Database D. Multiple pipelines, each serving a different business flow. We spun up Kafka instances on VMs, managed with Docker Compose as needed. They worked. We moved on. Over time, the seams started showing. Every change required SSH and port-forwarding. Our Kafka instances had no URL. You accessed them via localhost. To make any change, whether adding a source connector, adding a sink, or updating a configuration, you had to SSH into the server, port-forward to get local access, and make the change manually. Every time. In every instance. With Docker Compose, downtime was always one command away. Managing Kafka with Docker Compose works until it doesn't. Any update or fix that required restarting the stack meant running docker compose restart and hoping the restart went cleanly. For a data pipeline, that's not a comfortable position to be in. Versions were going stale, and updates were per instance. Kafka evolves. A new version was released. But because each instance was a separate VM managed independently, upgrading meant going into each one individually. To make it worse, our instances weren't even consistent. Some ran on ZooKeeper; some had adopted Raft as Kafka evolved. Different consensus models, different operational behaviour, different things to know before touching anything. Provisioning a new pipeline meant reinventing the wheel. No standard process. No shared configuration. Every new CDC pipeline was a fresh set of decisions: which version, which consensus model, which configuration. Knowledge of how each cluster worked lived with whoever set it up. Monitoring was manual. We used JMX Exporter to pull metrics and eventually shipped logs to Loki. These helped. But they were layers on top of an operational model that still required you to SSH into the right instance to actually diagnose anything. There was no unified view of the fleet. None of this was broken in a way that triggered an alarm. It was just slow, manual, and increasingly inconsistent as the fleet grew. We were spending engineering time on operations that should have been automatic. What Is Strimzi? Read the full blog post here Top comments (0)
Comments
No comments yet. Start the discussion.