Kemin’s IoT platform had grown to 17–18 services across five environments. Services could crash and recover before anyone noticed, and the logs needed to investigate disappeared with the restart. Often a customer was the first to report a problem.
FlowFactor built the stack in Kemin’s own Azure environment with Prometheus, Grafana, Loki and AKS. Logs are kept centrally, alerts arrive in Teams, and almost everything is defined as code so the team can extend it without us.
Visibility instead of guesswork
The dashboards surfaced a service that had crashed more than 100 times in a single week without anyone knowing. Kemin now sees memory, CPU, response times and data flows across every environment, and keeps developing the dashboards on its own.
Kemin is a global ingredient manufacturer. Family-owned since 1961, it makes more than 500 science-based specialty ingredients for animal and human nutrition, pet food and food safety, with customers in over 120 countries.
Alongside that business, Kemin runs an IoT platform that captures data from customer systems and turns it into insights.
Over five years, that platform had grown to 17–18 different services deployed across three production environments, a development environment and a test environment.
But Kemin had very little visibility into what was actually happening while it was running.
Limited visibility across a growing cloud platform
The production platform runs on Cumulocity, a third-party cloud environment. Kemin can access its own services and their logs, but has limited visibility into the infrastructure around them.
There was another practical problem too: when a service restarted, its logs disappeared with it.
Often, the first indication that something had gone wrong was someone noticing it by chance.
Or a customer getting in touch.
“Before, we only really knew something was wrong when we happened to be on the platform ourselves, or when a customer told us.”
- Peter Eraly, Digital Business Analyst
And these weren’t always obvious outages. A service could restart within 30 to 40 seconds, data might briefly stop flowing, graphs would freeze, or a message queue could fail somewhere in between systems.
By the time someone noticed, the service could already be running again, while the logs needed to investigate were gone.
Bringing in observability expertise
Kemin has front-end and back-end developers, but nobody specialised in monitoring and the infrastructure behind it.
The team had experimented with Grafana, but connecting multiple environments, centralising logs and metrics, and setting up reliable alerting required expertise it didn’t have internally.
Hiring someone permanently for that specific need didn’t make much sense either. Kemin wanted the initial observability layer built properly, with enough knowledge transfer to make smaller changes itself afterwards.
That’s where FlowFactor came in.
And for Peter, the way the team worked mattered as much as the technical expertise.
“We’ve had a lot of consultants come through. This was definitely one of the smoothest collaborations. They needed very few words to get something done.”
- Peter Eraly, Digital Business Analyst, Kemin
Kemin is a small team that works through conversation rather than extensively written requirements. After an initial meeting and occasional check-ins, FlowFactor’s engineers were able to work largely independently.
Peter particularly appreciated that they could read between the lines: understand not just what Kemin wanted to see, but why, and bring ideas of their own.
Building an observability stack in Azure
FlowFactor built the observability environment in Kemin’s own Azure environment using Prometheus, Grafana, Loki and Azure Kubernetes Service.
Logs are now collected centrally and remain searchable after a restart. Metrics feed Grafana dashboards that bring the different environments together. And when something crosses a defined threshold, the team receives an alert directly in Microsoft Teams.
Kemin can also trigger alerts from specific application logs. Instead of digging through millions of log lines, the team gets an indication of when something happened and can investigate from there.
Almost the entire environment is also defined as code. The underlying AKS infrastructure is provisioned through Terraform, while Prometheus, Grafana and Loki are deployed through pipelines and configured in YAML.
Very little therefore depends on manual configuration. The environment can be deployed or updated consistently instead of relying on someone clicking through the setup by hand.
Detecting production issues that previously went unnoticed
Early on, the dashboards revealed something Kemin hadn’t been able to see before.
One service had crashed more than 100 times in a single week.
Nobody had realised it was happening. The service restarted within 30 or 40 seconds each time, making the individual outages easy to miss. Once the pattern became visible, the team could investigate.
The cause turned out to be relatively simple: far too few resources and far too much memory usage.
Kemin can now monitor memory usage, CPU, response times and the internal behaviour of its services, helping the team tune them based on what actually happens in production.
Observability also gave them better visibility into the data flowing through the platform.
For one service processing MQTT messages, Kemin had been exploring different ways to detect when messages weren’t being processed correctly. FlowFactor proposed a simpler approach based on comparing incoming and outgoing messages.
It is exactly the kind of practical input Peter valued in the collaboration: helping the team simplify how they monitor and understand the system.
Knowledge transfer without creating dependency
The goal was never for FlowFactor to remain the only team that understood the setup.
Kemin’s developers can add graphs, adjust parameters and extend existing dashboards themselves. For larger changes, they can still bring FlowFactor back in when the expertise is useful.
Peter estimates the dashboards now contain three to four times as much information as Kemin originally asked for and the team continues to use them.
The collaboration also held up when the FlowFactor team changed during the project. Kemin didn’t notice a slowdown in delivery. Behind the scenes, other FlowFactor engineers could step in with additional expertise where needed.
That matters because the result wasn’t just a monitoring project that worked at handover.
It became something Kemin could keep using, adapting and building on.
Before, a service could fail and recover without anyone knowing.
Now Kemin receives alerts, retains the logs needed to investigate issues and has visibility into how its services and data flows behave across environments.
Running a platform you can’t fully see? That’s exactly the kind of problem we like to solve. Let’s talk it through.
