AI in DevOps: where it belongs in your CI/CD pipeline, and where it doesn't
The instruction from the top is often short: use AI, it will make things better. For AI in DevOps, that usually ends with an agent somewhere in the middle of the CI/CD pipeline. The pipeline gets slower, the token bill grows, and nobody can point to what improved.

So where does AI actually belong in the software development lifecycle (SDLC), between code and production? Our answer fits in one line. AI won't make your pipelines faster. It makes you faster when your pipelines fail.
Four questions before you add AI anywhere
Before any agent goes into your delivery process, ask these four questions.
Do you need something deterministic? If the answer is yes, put it in a script and stop there. An agent is not deterministic, so it is the wrong tool when you need the same output every time. Put one in the middle of a predictable process and you also make troubleshooting much harder.
Does it need judgement over a lot of input? Then it gets interesting. Think of large volumes of logs, or legacy configuration that someone has to reason about. That is where an agent adds value.
Is it on the critical path? Then don't. AI gives advice, and advice is fine when it is wrong. A blocked merge request because of a wrong assumption is not. The cost there is not tokens. It is time. As platform engineers, our goal is to get feedback to developers as fast as possible and keep releases moving.
What does it cost? Tokens are the obvious part. Developer time is the part people forget. Developers who wait longer for a pipeline cost money too.
With those questions in mind, look at AI in CI/CD in three layers: before the pipeline, inside it, and after it. The biggest value sits before and after. The hype puts agents inside the pipeline itself, and that is where they add the least.
Before the commit: spend the tokens once
This is where AI earns its place. It is good at four things here: generating, migrating, validating and documenting.
Generating speaks for itself: scaffolding for pipelines and larger structures. Migrating means turning legacy into something maintainable, for example converting Jenkins pipelines to GitHub Actions, or shell scripts into Ansible so they can be automated.
Validation is the one teams often skip or neglect. Variable validation in your infrastructure code almost never gets written beyond "optional" or "default". AI makes it cheap to go further. For example: this variable is optional, but it becomes required when that other one is missing. The same goes for assertions in Ansible.
Documentation is the easiest win. Nobody wants to write it. Let the model do the first draft.
The rule that ties this together: spend the tokens once. Let AI write the script, then run that script deterministically from then on. There is no reason to have an agent generate and execute the same script on every run.
Why does that matter so much?
Cost. An agent that regenerates the script on every run makes you pay for it on every run. Generate it once and you pay once.
Predictability. A pipeline goes from A to B with the same steps, every time. An agent does not. Give it the same task twice and you can get two different results.
Troubleshooting and audit. A script can be read, reviewed and versioned. When something breaks, you know exactly what ran. When an agent improvises a slightly different approach each time, finding the cause of a failure gets a lot harder.
And review everything. Infrastructure code deserves the same testing and validation as application code. A YAML file that passes a syntax check has not been validated.
Inside the pipeline: advice, never the gate
This is the layer with the least value, so keep AI on the side. Your AI must never become your quality gate. It is supposed to be a bit creative, which also means it is allowed to be wrong. Let your linter and static analysis keep doing the blocking: they are fast and they are free. AI runs next to them, in parallel, and only advises.
Within those limits, there are a few places where it does help.
Use it for triage, not scanning. Security scans often produce failures or huge outputs. An agent that summarises them is useful. An agent that replaces the scan burns tokens and takes longer than the tools that already carry that load. If you want an AI review of your codebase, a nightly run after the day's releases is a better fit than a run on every change.
Prepare, don't merge. Platforms already open a pull request when a new dependency version appears. That is the right pattern. An agent can prepare the change, but it never merges it on its own.
Run it only when something fails. If you don't want AI in the pipeline at all, this is the light version: trigger it only on a failed build or compile. By the time someone opens the logs, a first triage is waiting. That is exactly where the time goes.
After the pipeline: make yourself faster
The second layer with real value is everything that happens around your runs.
Triage of incidents and log aggregation is the obvious one: AI that pulls the relevant signals together and points you in the right direction before you start digging.
Larger maintenance runs are another. Our team updates specific dependencies across repositories every quarter. Every update triggers builds. Afterwards, someone has to work out what succeeded, what failed and what still needs follow-up. Sometimes a script can produce that report.
When the information is spread across several systems, an agent that follows the run and summarises the outcome is a good fit.
And evaluation: rightsizing and cost efficiency. Let AI analyse where workloads get too many or too few resources.
Guardrails that keep it safe
A few rules we keep regardless of where AI sits.
Read-only, for as long as you can. If you give an agent write access, you also take responsibility for everything it changes. You don't want that on your conscience.
Split reading from doing. A single job that reads all input and also writes has rights on the whole repository. If its credentials leak, an attacker can reach every repo. A safer setup: one agent reads and proposes, and a separate apply job does the change with a just-in-time token that is valid for fifteen minutes. A leaked token is worthless, and splitting input from apply also limits what prompt injection can do.
Keep a human in the loop and an audit trail. The advice is the deliverable. Make sure you can always see what the agent is doing and why.
Put a budget kill switch on it. Use an AI gateway or a token budget. If costs shoot up after day one or two, it switches off automatically.
Never depend on your provider. If the connection to your model provider drops, your pipelines and releases must keep running.
Know where your code goes. If you don't run the models yourself, check which model you use and where your code and data end up. Some input you may not want to share.
Where to start
Measure first. What does your delivery cycle look like today, and which step is the slowest? Very often it is the time spent fixing things when they go wrong.
Pick that step. Add AI outside the critical path, with a hard kill switch. Then measure again. Did it actually help, did nothing change, or did it get worse? Then take the next slowest step and repeat.
AI won't make your pipelines faster. It can make you faster when they fail. Your pipeline should keep running either way.
If you want a second opinion on where AI in DevOps fits your own delivery process, that is a good conversation to start with an audit. If you are setting up or rethinking a pipeline, we also wrote about building a CI/CD pipeline and the choices no vendor will tell you about. Building that kind of platform is what our platform engineering team does every day.
Kubernetes & OpenShift
Application Gateway for Containers in production: lessons from an ingress-nginx migration



