Copy logo as SVG
Copy wordmark as SVG
Download brand assets
Brand guidelines
Blog Vision

How we run Tuist with AI agents, Kubernetes, and Elixir

Pedro Piñera

What an exciting time to be a builder. For the first time in a long time, we are even questioning whether GitHub will remain the dominant go-to option for hosting code, or whether pull requests will still be needed. No one knows what the future might look like, but everyone is trying, and we find it exciting.

At Tuist we’ve been open minded about these things. Trying new ideas, challenging assumptions that emerged in a context that has just changed, and overall, trying to get the most out of it so that we can maximize the value per person in the company. Someone recently described us as “AI-pilled.” We take it as a compliment. The way we work today has nothing to do with how we used to work months ago, and might be far from how things will look a year from now. However, what we’ve learned over these past months, and what we believe will hold true, is that we built convenient abstractions that made sense at the time for humans, but that might hinder the effectiveness of workflows driven by agents when they interact with our organization. In this blog post I’d like to share some examples of areas of Tuist where this has manifested, and what happened when we gave agents access to the systems underneath those abstractions.

Kubernetes infrastructure that AI agents can inspect

When we started working on Tuist, we did what most builders did at the time, and what many still do: use a platform as a service. Many of them are just layers on top of cloud providers like AWS, conceptually compressing the intricacies of dealing with VMs, load balancers, ingress proxies, and other sorts of infrastructure-related pieces. Think about it: I built this software service and want to get it to production quickly, so ideally my deploy is as easy as a git push, as Heroku pioneered at the time and others like Fly.io followed. I like to think of these layers as languages. At the bottom we have bare infrastructure resources, like servers. Cloud providers abstract them with another language that treats them as resources to rent out, ensuring they can maximize the usage of the underlying resources. That’s still not high level enough for most people to use, so a higher-level language exists, letting you talk in terms of what you work with, for example “apps” which you deploy.

For a long time that’s how we ran Tuist. We even switched providers, but as Tuist grew beyond a single service, and we needed low-level control to optimize latency and resource usage, we started considering going one level lower. For people like us who had never used Kubernetes, this felt a bit daunting. We even played a bit with Kamal, but it took a few conversations with agents to realize that LLMs speak the language of Kubernetes and infrastructure. It felt eye opening. We went from agents not speaking the language of those platform products, and requiring MCP as an interpreting layer, to agents speaking an established infrastructure language. That would give us the flexibility we were looking for, but also the flexibility to port our infrastructure stack to another provider, or get agents to debug the state of the infrastructure at any time using the runtime interfaces provided by Kubernetes. It felt empowering.

Many still see Kubernetes as that one tool just for big enterprises. There’s some truth to it. Kubernetes is accidentally complex, and agents dynamically compress that complexity at runtime, so you can learn about what you need to achieve your goals and get your stack deployed to the bare resources. That has a side benefit of lowering costs.

There’s also something beautiful about the Kubernetes language, and it’s the fact that it’s extensible. We sometimes joke about the fact that we’d model many things as state to be reconciled. But it’s true... we ended up incorporating things into that deployment model that used to live in separate continuous delivery pipelines. An example is Cloudflare configuration, which we applied using the Wrangler CLI from a delivery pipeline, or in some cases, manually updated through their UI. Now it’s just state, versioned in the Git repository, with our custom controller reconciling that state on merge. We can have a conversation about building these kinds of extensions, about the modelling of the state, about how the state is being reconciled, about the history of changes to that state, and about the state right now. It’s beautiful. We have an interface to understand the system, including details that a platform might abstract away to make things convenient. Often you need to understand your system at a deeper level, and the last thing you want is to lack that access and have to resort to support tickets. We suffered through that with some outbound network issues that eventually led us to move to another provider.

The second thing that became clear through all of this is that the UI is very secondary. Runtime access to the infrastructure can be completely headless because agents don’t need your UI. We just need an interface to access it, with the right mechanisms in place to do so safely and tools to elevate privileges in rare scenarios, and a language that agents speak. Kubernetes is one of those languages. All of our work to iterate on the infrastructure happens through conversations with agents to either adjust the declaration or extend the language, or conversations to understand a particular state of the infrastructure, for example when some piece of it is misbehaving. Kubernetes is not just how we deploy the servers, but also how we get our runners into a desired state to run our customers’ CI builds, or configure Cloudflare to stop crawlers from exhausting our resources while crawling public projects.

Adopting a common infrastructure language has another side benefit: our on-premises customers can use our charts to deploy the Tuist stack to their own infrastructure. Otherwise, we'd end up using a different deployment configuration from our customers, making it trickier to support them without dogfooding the thing we'd provide them with.

Debugging production with Elixir and Erlang

Sometimes, agents need to go beyond the infrastructure and dive into the pieces that are running in those pods. I've always said it, and I'll keep saying it. Elixir and the Erlang runtime make up one of the most productive combos that I've ever seen, and I'd recommend them to anyone. When we chose Elixir, people warned us that it would be a problem because there weren't that many developers who knew the language. It's funny how things have changed. It wasn't much of an issue for us at the time, and with agents, it's even less of an issue today. Sometimes, the server is misbehaving, or one of the workflows is not working as expected. Most fixes start with a reproduction, but the local environment might be too different from production to debug the issue there. In scenarios like this, it comes in handy to tap into that production system, obviously being careful about it, to inspect what's going on. This is like doing surgery after an accident, where you need to open things up, understand what's happening, and then come up with a fix that you can either apply there or translate into a local fix that goes through continuous deployment.

This talk does a better job of explaining it, but in short, the Erlang VM can be inspected at runtime. Agents can open a console against the running application, with access to all the processes running there and their symbols, and interact with them. Not only that, but there are runtime inspection interfaces. For example, you can sort the processes by their memory or CPU usage, or locate a process by an identifier. And not only that, the underlying scheduler is fair, which means that even if it's busy handling a ton of traffic coming from clients, it leaves room for you (or your agent) to run one of those inspection sessions and understand what's going on there.

Agents speak the language of Elixir, and especially Erlang, quite well, since Erlang has been around for a long, long time. Compare this to, let's say, a JS serverless function running on a platform that doesn't let you inspect its runtime. It might be misbehaving or consuming unnecessary memory, and there's no way for you to tap into that environment because it was never designed for this kind of debugging. I believe more and more that having access to runtimes at every level can make a huge difference between iterating fast and iterating insanely fast. We want to be on the latter side, and that guides how we choose technologies. The same thing that made Kubernetes empowering shows up here: agents can speak the language and reach the running system.

Observability with Grafana and AI agents

Infrastructure ensures our software has a place to run, and Erlang's runtime runs our code while giving us an interface to inspect what's going on when needed. However, knowing when something needs our attention is something we need to invest in too. Once again, we want an observability runtime that agents can inspect. Turns out we've had standards around observability for quite some time: Prometheus metrics, logs, and OpenTelemetry metrics. Agents understand them. They can also query them... We went ahead with hosted Grafana, which would take care of storing all that observability data and provide us with a couple of pieces, among others, that we'd start with: alerts, which we could wire into an incident response management system, and dashboards.

With dashboards, something really interesting happened. I remember some time ago listening to a podcast where the interviewee said that there'll be a time when we'll deploy things so fast that we won't have time to look at dashboards, and therefore dashboards won't be needed, at least in the static form that we know today. If an agent can monitor the data as it comes through, why would you want a human looking at a dashboard, especially when they might be very ineffective at drawing conclusions from it? Because that's the purpose of a dashboard, no? And let's admit it, doing so at scale is something we are not good at. Unless you like looking at a dashboard for the sake of looking at the dashboard, as was once the case in a company I worked for, where a director of engineering would spend hours looking at them as a sign of pride. So we have dashboards, but we barely look at them. Instead, we use Grafana in a very headless way. First, we define alerts, a lot of them. From alerts that tell us if queries or page loads are slow, because we want Tuist to feel snappy, to alerts that tell us if there are unusual error responses in our system. Some are routed to our Slack, others directly to the incident response system because we are confident they are worth escalating.

The moment shit hits the fan, we trigger an agentic session with access to query all that data through the MCP interface in our operations platform. It acts as a proxy so people can access it without authenticating separately against Grafana's own protocol interface. That means if something happens, an agent can use that interface to understand the system at every level, using the local cluster definition in the repository as a reference, and going as deep as opening a console against the Erlang runtime. The infrastructure, runtime, and observability data become pieces of the same investigation. Our usage of Grafana has become very headless, and it has in fact motivated me to tinker with Pulso, a stateless, object-storage-backed solution for storing and querying observability data and sending alerts based on it. It's built in Elixir, leaning on Rust for the performance-critical pieces. It's still a work in progress, but the idea is to put it in production soon and potentially replace Grafana with something that we can operate, scale, and improve as we need.

Company operations with Atlas and MCP

I talked about it here, but the same idea of having a runtime that agents can interface with applies outside of engineering too. I got obsessed with engineering the organization. In part, because I want to provide the best value to our users and want our organization to stand out among our competitors. In part, because we can't afford to have many specialized roles, and I honestly believe many are not needed. And in part, because I want to have space to do more building. When I started moving pieces of context to what we called Atlas, connecting pieces of information, and then using chats with LLMs to have conversations about areas of the organization, it felt surreal. Just yesterday, I took all of this one step further and connected it to Grok Bots, something that OpenAI has released recently as Dots, and it felt surreal again. I created virtual roles, for example someone responsible for finances or support, and gave them clear details around how we operate financially, or how we'd like to provide support to our users. I played a bit with it, using the mode where it feels like having a call with someone, and I was in complete awe. This is how I imagine the future being. Not trying to make everything autonomous, but having all the right information in place so that you can have conversations about it with agents and with your colleagues, make the best decisions, and then act on them. We don't see Tuist running autonomously. However, we do see many processes being triggered by signals, including signals that we derive with the help of agents from all of the context we've collected over the months.

Our intelligence runtime is Atlas, and we access it through an MCP interface. Just yesterday, I requested a tax certificate from the German tax authorities, which involves sending them a letter by post, and it was just a prompt from my phone, on the go. Isn't that crazy?

Standards give AI agents a shared language

Having a context runtime, and an interface to it, in every domain of the company can make a huge difference between being efficient and feeling sluggish. This might require revisiting many past decisions. In our case, we chose Elixir without seeing this coming, but even then we started to appreciate the value of access to the runtime. Agents just augmented it. Our move to Kubernetes wasn't motivated by that either. We were morphing into an infrastructure company, so we needed a way to describe it. As a side effect, our agents found themselves with an interface that we could lean on to iterate further on the infrastructure and debug its state at any time. And the work on operations also happened organically, this time truly motivated by building that kind of intelligence runtime. Across all of these areas, the same idea keeps coming back: we must own the context, and we must connect it in a way that lets us make great decisions quickly, without having to build departments or raise money around promises to hire people for those roles. We want to stay intentionally small, maximizing the value we capture per person in the company, and this guides a lot of the decisions that we make.

For me, all of this comes down to having someone to talk to about the systems we build and the company we run. Someone who has access to the context, can look into what's happening, and help us think through what to do next. For that conversation to be useful, our systems need to speak a language agents understand. And surprise, surprise, that's what standards are all about. Bet on HTML, on CSS, on JS. Bet on OpenTelemetry or Prometheus metrics. Bet on Kubernetes. Bet on peeling back convenient abstractions that get in the way of those conversations, and on owning and connecting your data so that you and your agents have something meaningful to talk about.

What are you waiting for?

La IA se mueve rápido. Tus builds también deberían.

Comienza Habla con nosotros