Photo
Photo
Why Big Companies Still Can’t Trust AI | Atin Sanyal, Galileo

381 / July 10, 2026

Why Big Companies Still Can’t Trust AI | Atin Sanyal, Galileo

70 minutes

381 / July 10, 2026

Why Big Companies Still Can’t Trust AI | Atin Sanyal, Galileo

70 minutes
Listen on

About the Episode

Your AI agent aced every test. So why is it quietly failing in production?

That gap between looking perfect in testing and staying trustworthy at scale is the problem Galileo was built to solve. Its evaluation and observability platform now helps Reddit, Airbnb, Procter & Gamble, Comcast and six Fortune 50 companies ship AI they can actually trust.

In this episode, Atindriyo Sanyal joins me to trace Galileo’s journey from the earliest days to its recent acquisition by Cisco.

Before founding Galileo, Atin spent a decade at Apple and Uber, where he helped build Michelangelo, Uber’s AI platform that’s still a blueprint for modern MLOps.
We get into:

– Whether small models actually beat large ones for AI evaluation
– Why bad data is one of AI’s most underrated risks
– How evaluation and observability became the trust layer for enterprise AI
– The real trade-offs between cost, latency, and quality
– What it takes to build AI enterprises can rely on

If you’re shipping AI products or running agents in production, this one’s for you.

Watch all other episodes on The Neon Podcast – Neon

Or view it on our YouTube Channel at The Neon Show – YouTube

Siddhartha Ahluwalia 0:52

Hi, this is Siddhartha Ahluwalia, your host at Neon Show and managing partner at Neon Fund, a fund that invests in some of the best enterprise AI companies, start from India and building globally like Atomic Works, Spotdraft, CloudSEK. Today I have with me Atin Sanyal. Atin, welcome on the Neon Show.

Atin Sanyal 1:08

Thank you so much for having me. I’m super excited to chat.

Siddhartha Ahluwalia 1:12

Atin, your journey has been so inspiring. You have been a researcher, came from India, came to do your master’s here. Then you were part of the Michelangelo team at Uber. And then you started Galileo. So there is very similarity with Michelangelo, Galileo. Did the name for Galileo got inspired from your project at Uber?

Atin Sanyal 1:32

That’s a great question. So the story goes that I was at Uber working on Michelangelo, and I had met my co-founder, Vikram. Vikram was at Google, and he was working on like finance systems like Google Pay, and then did a bunch of AI work. And he sent me this WhatsApp of a diagram he drew on a whiteboard of a dashboard. So we saw charts, etc. And at the top, he just wrote Galileo. And I asked him, what is Galileo? Did you just already come up with the name of the company? He’s like, no, I just threw in some similar medieval character like Michelangelo. And I was like, I love the name. So it kind of became a thing. But turned out, Michelangelo, the real person died the same year that Galileo was born. And I found that out two years later. So funny coincidence.

Siddhartha Ahluwalia 2:36

And which was your first job after UCLA?

Atin Sanyal 2:38

So I joined a very early post acquisition Siri team at Apple. And the team was about 30-35 people, still the founders were there. And I ended up, it was a long story to join that team. I ended up working there for a good number of years, about five years, where I worked on the early knowledge graphs that Siri was built on. And it was kind of in the era of early deep learning models were being experimented. It was still very expensive to run it in production. But I kind of saw the whole arc of how do you build knowledge graphs and language learning systems without these models and then saw models kind of become a thing.

Siddhartha Ahluwalia 3:21

Got it. And then how did Uber happen for you?

Atin Sanyal 3:26

So Uber was an interesting opportunity that came to my plate. At Apple, I had worked towards the end of my stint at Apple, I was working on some streaming machine learning based systems, more from an infrastructure standpoint. And Uber was starting out this team. It was supposed to be the foundational AI team. And they required engineers who had worked on scaling streaming systems, because streaming ML was very new. This is back in 2016, 17. So I decided to talk to them. And they were very interested in getting me to solve that particular problem. And then over time, I kind of became one of the leads and architects of Uber’s entire AI fleet. So Michelangelo, for those who don’t know, is Uber’s AI platform, which hosts thousands of models across demand pricing, ETA prediction, any AI feature that you’re using across any Uber’s product today is powered by that platform. But back then, it was very zero to one, we barely had any models running and Uber’s machine learning fabric was just a bunch of data scientists, training their own models. So it seemed like a really exciting opportunity. And I also loved that Uber was one of the first systems that truly worked on very mission critical machine learning, because you’re putting two people inside an enclosed box. And there’s safety issues and a lot of new challenges that the world had not seen before the marketplace economy had come. So it seemed like a very exciting technical challenge. But over time, Michelangelo kind of became a brand in the AI infraspace. Even today, I get emails and emails from folks talking about how they’re still using the Michelangelo blueprint to build their internal AI system. And just the name is enough, at least in the community. I feel it’s the most underrated team because it was a team of absolute superstars, some of the best in the world, AI experts. And it gave me the foundations on how to think in systems and how to really build world class software. So I owe a lot to my experience at Uber. And it was kind of the foundations of not only me personally starting Galileo, but also learning how observability and evaluations is such a big problem in machine learning. So even though the workflows have changed today, the fundamentals still remain the same. The problems that people are facing on how do I know whether the output of a machine learning or an AI system is good or bad, they were the same challenges we were trying to solve. But the mission criticality of it is higher than ever. Because I think it’s the first time in human history that AI has become the face of the product. A user is directly talking to an AI and working with an AI. So every output of an AI system is mission critical. And with agents doing actions and making critical decisions, the stakes cannot be higher.

Siddhartha Ahluwalia 6:43

And how did you and Vikram zero in on the idea of Galileo?

Atin Sanyal 6:48

That was a long story. So we were discussing on what to build. Vikram was certainly very gung ho about starting a company. And I was deeply in the AI space. So I knew that AI is the future. And I was very excited to build something in that. He also saw the same sort of vision. We initially started off working on some data infrastructure for machine learning systems. And while we were ideating and building, we realized that trust is a big problem in machine learning. And it is unsolved. So we started building a trust layer of every machine learning workflow. How do you quantify whether an output of an ML model is good or bad. And one thing led to another and I created a prototype. We met a professor at Stanford, Chris Ray, who we knew through one of Vikram’s friends, who had recently sold a company that Chris Ray was involved in. So I showed him this little demo of a Jupyter Notebook, where you the user publishes the model and before the model gets productionized, it would do these bunch of checks. And Chris Ray saw it and he’s like, this is a great idea. We will do it for language models. And my only question to Chris was what the hell is a language model? Right? Because LLMs were not a word back then. And of course, I knew NLP systems and stuff. But this idea of running language models at scale was so alien, and so new.

Siddhartha Ahluwalia 8:21

This is 2021.

Atin Sanyal 8:22

This is early 2021. So late 2020. But it seemed like unstructured data AI is the future. And I realized that because even in my time at Uber, we were slowly productionizing more and more deep learning models. I still remember a moment at Uber, where we productionize the first deep learning model. And we saw that it’s kind of giving the same results as an XGBoost classifier, but it’s so prohibitively expensive with GPUs. And so we realized at least for some classical use cases, like demand pricing and those kind of tasks, it’s not really needed. It’s like using a fancy pen when a pencil would work. But over time, the cost of running these models in production went lower and lower. And that was kind of my main aha moment, where it was clear to me that the next 10 years, 99% of the world’s data does not live in tables in cells. It takes a lot of machinery to get data into a clean format. And for me, it was a personal realization because we had worked on the feature store at Uber. And we had built the world’s first feature store. And we literally… A feature store for the audience, especially, is… Think of it as a database of data that’s ready for models. The question is, how do you get data that’s ready for models? It’s a long, arduous process of doing ETLs on raw data, which is very messy, very raw. And you set up these pipelines, which converts that raw data into very specific ML data, which is ready for a model to consume. Features are essentially the data points that go into a model. An example of a feature, say in Uber’s case, would be the number of rides taken by a user in one week. Now, that value is changing as the days go by. So there’s this whole rolling aggregation that you have to do to constantly keep that feature updated. And because a stale feature, if it goes into a model and a model makes a decision on that stale data point, it’s the wrong decision. So a feature store is essentially this whole data management system, which abstracts all this away. You simply define a feature and then you give it some configurations. And it just makes it ready for you. And you can just consume it via API. So this idea of building this data ready for machine learning systems really blew up. But I realized that the cost of bad data in AI systems is catastrophic. It leads to wrong decisions. And those wrong decisions percolate to downstream systems. And you end up with an answer that the user sees that makes absolutely no sense.

Siddhartha Ahluwalia 11:25

Got it. And let’s say when you heard in about 2021 about language models, can you share in now that they are popular, but to our audience, what does language model mean back then? How do you understand to a layman?

Atin Sanyal 11:40

Absolutely. So funny enough, the definition of a language model hasn’t really changed much over the years. So to dumb it down, a language model is a machine learning model. And a machine learning is of course, when I say a machine learning model, it is essentially a deep neural net. So all these models like OpenAI, OpenAI’s models, Anthropx models, if you shine a magnifying glass on them, they’re essentially a deep neural network, which has very specific architectures behind the scenes. But what they really are, are any deep neural network is essentially a memory machine. And the whole process of training a model is making the model remember patterns in the data. But data to a model, it’s just ones and zeros. So you can feed it language, you can feed it images, and it makes no difference to a model. It will just remember the patterns and make predictions based on that. So a language model, in the more modern sense, is what used to be called as a seek to seek model, which is the output of this neural net is essentially one token generation after another in sequence.

Siddhartha Ahluwalia 12:57

How do you define a token?

Atin Sanyal 12:59

Okay, a token is a, for simplicity, think of it as a word. So if you have a sentence, say, how are you, the three tokens or four tokens in them are how are you? And then the question mark. That’s the simple definition. In reality, what a model does is it breaks down these words into sub tokens and multiple tokens. But that’s, you know, that’s essentially machine learning language, which is abstracted away from the user. But that’s the reason why you see any LM basically spit out one token at a time. And that’s why you see the streaming. It’s essentially at each decision at each token, it’s making a choice of what token to spit out next. And that choice is essentially a probability distribution of all of vocabulary. So think of the entire English dictionary. That’s whatever, half a million to a million tokens. So at each decision, the model is choosing what’s the next best word, but based on the previous word, using techniques like attention and a few other concepts. But it’s really just choosing one word and the choosing the next best word to say after each previous word.

Siddhartha Ahluwalia 14:17

So where did this insight come from that? Hey, obviously, you were at my current angle that observability for these large language models and AI in production could be the future.

Atin Sanyal 14:30

So I realized that AI has always had a measurement problem. And this goes far before language models were even a thing. Even simple models like XGBoost decision tree classifiers, it was always hard to quantify good or bad. The output of a model, which is a decision, how do you know whether it’s good or not? Except in the previous world, when we use simpler models, the input to them was these features that we were talking about. And often, it’s easy to do certain basic correlations to really dig into why a model chose a particular decision. The whole ballgame is different with language models, because language models are these token spewing memory machines, which is essentially, there’s so much randomness and probability in them. It’s very hard to say that this was the final answer based on that. Why did the model say what it did? It’s a hard problem. And that makes observability of these newer systems even harder. But that was the original problem that we were chasing. Can we dive deeper into these models and really understand why a model chooses the next word to say? And a lot of our initial algorithms that we had built in the early days of Galileo was essentially using math and statistics to quantify uncertainty. So, uncertainty is another technical word, but it’s really, in simple words, it’s how sure was the model that the next word is the right thing to say? And there’s some signals you can get from these models. So, we took the seed of that idea and built some algorithms to give you how certain was the model in its answer. And from there, we’ve done other more advanced methods to really quantify hallucinations and some of the more advanced issues we see. But it really comes down to this whole idea of uncertainty. How do you quantify uncertainty in these models?

Siddhartha Ahluwalia 16:34

Who are your first customers? And what’s the first use case that you solved for the first 10 customers?

Atin Sanyal 16:39

Yeah. So, given we started before LLMs were a thing, a lot of the workflows, especially in the enterprise, was driven around fine-tuning some of the smaller language models. So, for those who’ve only learned of language models after ChatGPT, they wouldn’t know about the erstwhile models, but these models were already there. In fact, even before GPT 3.5, you had GPT 2 and GPT 1, which were essentially similar architectures as ChatGPT, the modern ChatGPT, except the size was much lower and smaller. And in the AI world, we measure the size of the models through what we call the number of parameters. So, if you ever hear people say, I have a 500 billion parameter model, they are basically saying they have a really large model. So, coming back to the workflows, folks were working with much smaller models back then, typically in the 100 to 400 million parameter range. So, very small size. It’s so small that in theory, you could run them on CPUs. They would be slow because it’s sequential. But the GPU infrastructure also needed to serve these models was much less resource intensive. And the workflows back then were primarily around contact center AI and entity detection. And these are what we call tasks. So, entity linking, entity detection, classification, these were standard language modeling tasks that enterprises were doing. And our first product was essentially trying to make the model decisioning of these tasks much higher in quality. And we realized that the main determinant of that is the data that you feed into the model. Because the model is just an equation. There’s not much you can tune to it except a few hyperparameters. But that’s not the real fuel that you need to improve the quality of the model. And our first customers were mostly innovative startups, and mostly in the contact center AI space or similar thereabouts. I remember our first big customer was a very large bank in the US. And they had many contact center AI teams. They were also building their native checkbox. But this was again in the previous era before LLMs. And all those teams have now shifted to LLMs. And they continue to be our customers because we moved with them.

Siddhartha Ahluwalia 19:13

And how did the product evolve since then?

Atin Sanyal 19:17

We’ve had a very interesting journey with our product given we’ve seen the entire arc of language models for the last six years. We originally had this product, which essentially was a data scientist tool. And data scientists are folks who train and evaluate models. And they really understand the depths of these model architectures. And the main thing they actually work on, though, is curating data. Because at the end of the day, you’re not really a scientist sitting building new model architectures. Most data scientists are essentially data curators is our realization. So data was the biggest leverage, biggest bottleneck. And we built a platform that would essentially hook into these model training jobs and detect issues in your data using the same statistical modeling techniques. And that became quite a popular tool, which many of the data scientists in our early customers really loved because they saw 60 to 70% improvements in their model quality in a single shot simply by using our product versus it could take…

Siddhartha Ahluwalia 20:32

Because your product could produce better quality of data?

Atin Sanyal 20:36

No. So, well, yes. But the main thing our product did was detect issues in the user’s data, including label issues, including data issues in a single shot. So you could essentially just grab all the data that you have, pass it through our system, and we would tell you exactly what the high quality data you should train your models with. So it became this pass-through filter, which saved weeks and months of work that data scientists were originally doing.

Siddhartha Ahluwalia 21:08

But what were data scientists at these enterprises back in 21, 22, working on such cutting edge problems? They were trying to train their models using clean data?

Atin Sanyal 21:19

That’s a good question. There were some cutting edge teams who were working on not the generative type of models. That’s a whole different part of our story. And we did get into generative model even before ChatGPT, because there were far and few teams who were doing these sort of pre-ChatGPT generative model use cases. But those use cases were far and few. The main use cases were kind of these models acted as behind the scenes decision makers. But the traditional chatbot systems usually had a business logic and code that would power them. And that code would use the decisions from the models, but the model’s output would never reach the user. And then there were maybe 1% of our customers were doing seek-to-seek modeling. Now, seek-to-seek modeling is the token generation modeling, which today is known as LLMs. But they were very far and few because this whole use case around, hey, I’m going to use a small model and fine tune it and make it spew out tokens to generate a sentence. The quality of the sentence generation was very poor. In fact, you can go to Hugging Face and try out some of these older models. And I assure you, for anyone who’s only worked with ChatGPT, they’ll be shocked. In 2020, I think at OpenAI, they had done a lot of experiments on having step function increases in the parameter size. That’s what led to these models actually constructing coherent sentences. But for the vast majority of the world, these generative models could barely talk. They were like a baby. They would say some random English sentences, which would not make sense. And there’s a lot of old videos on Twitter, you can see. But those were not our primary use cases. That’s why it got really exciting when ChatGPT happened for us because we’ve spent all these years working on language model research. How can we keep our research DNA but apply it to the newer cutoff models? And now we actually have real businesses who are very excited about using this technology. So it was around 2023 or late 2022, early 23 is when we are like, man, this could be a really big business.

Siddhartha Ahluwalia 23:43

So observability, you got this insight with the launch of ChatGPT in 2023, that it could become a very large business.

Atin Sanyal 23:55

That’s right.

Siddhartha Ahluwalia 23:57

But what about the enterprise reduction? How fast or slow were these enterprises from 2021 to 2026 early? If you have to lay out that auction curve?

Atin Sanyal 24:10

Yeah, that’s a great question. So I think my observation, just looking at our customers, our users, has been that there’s in many of the enterprises, especially FSI, banks, telecom, they have AI centers of excellence, and many follow this hub and spoke model, where you’d have one AI team that would build out the tools and the technologies for others to consume. And so that model still existed back then. It was just that the AI centers of excellence were filled with very niche skilled data scientists who likely had PhDs and were very good at unstructured modeling, because a lot of AI back then was essentially based on structured data. And while there were those teams, the use cases were very different, at least for us. And we only wanted to work on unstructured AI use cases, be it language, voice, images. So a lot of those teams were very specialized teams, often research labs that would hire PhDs only who have worked on this stuff. And it was a very niche skill. That, I think, led to teams moving slow in general. One, it’s hard to hire these people. And two, a lot of these people would want to join a Google or some big tech sort of fang-like company who would pay more. So there were all these challenges. Ever since ChatGPT, I’ve seen the adoption of AI in enterprises increase its velocity manifold. And the reason is because the niche skill has been commoditized. Now you have a black box. And the whole point is how do you build a software and infrastructure layer around this black box that can make it more predictable, you can control it better. That’s not a modeling problem. That’s a software engineering problem. So banks and telecom companies and healthcare companies, while their engineering and innovation has stepped up in its velocity, the main bottleneck now is compliance, security, and just general resilience. How do you build reliable AI? Which is the hot question. And that puts us in the eye of the storm as a company.

Siddhartha Ahluwalia 26:35

And LunaOne was your first model that you built. Why did you build a model in the first place?

Atin Sanyal 26:43

Yes. So our Luna story is fascinating. So for those who don’t know Luna, Luna is a small language model, which is specifically designed to solve the evaluation problem in AI. So how do you know whether an output is good or bad or an input has some PII or any security issues? So these are very specific tasks, which Luna is designed to solve for. So by default, the size of the Luna model is much smaller. It is orders of magnitude smaller than, say, a general LLM, which is out there. In fact, our latest Luna models, which are some of our largest, they typically range from 1 to 3 billion parameters, which is miniscule compared to some of the larger foundational models out there. Luna came into the picture when we realized that LLMs as judges, which was the traditional way that people were using to evaluate, and that has a whole history of it. They don’t scale. They don’t scale in production. They don’t allow you to do full scale observability, which means intercepting every single input and output of these AI systems. They simply choke. They are very expensive, and they’re highly unoptimized to solve for the latency and cost problem. So we took this problem and kind of came back to the drawing board saying that, hey, how do we distill all this intelligence and reasoning abilities of LLMs and bring it down to a much smaller particular set of tasks, essentially converting these writers, like these LLMs are basically token spewers and writers into calculators. That’s the difference between the main difference between Luna and a general LLM is that Luna is optimized for a very limited space problem, and it solves that really well versus LLMs are optimized for general reasoning and general generation capabilities.

Siddhartha Ahluwalia 28:49

Got it. So stepping back, one of the core areas for you also became during the course that latency has been a big issue, right? And for you working for the future also until now, how did you attack this problem of latency?

Atin Sanyal 29:10

We saw it firsthand that when you make a call to an LLM, it takes many seconds.

Siddhartha Ahluwalia 29:15

15 to 20 seconds.

Atin Sanyal 29:16

15 to 20 seconds, sometimes even higher if you’re in thinking mode, and you’re trying to…

Siddhartha Ahluwalia 29:21

Or deep research.

Atin Sanyal 29:22

Deep research takes, of course, minutes. And we knew that if enterprises and businesses are going to adopt this technology, they’re not going to sit and wait for 15, 20 seconds, especially for observability, where time is of the essence in observability. Latency is one of the top three things you solve when you build an observability system. The whole point of observability is detecting an issue right when it happens. And signaling it back to the user. So this is a solid problem in traditional software.

Siddhartha Ahluwalia 29:57

Yeah. SRE has been doing it for decades.

Atin Sanyal 30:00

SREs have been doing it for decades. The template is out there. There’s many unicorns and billion-dollar companies which have come out in the observability space.

Siddhartha Ahluwalia 30:10

Planck, one of them.

Atin Sanyal 30:11

Planck is one of them, yes. And so it has always been a mission-critical problem. But the key factor there is every observability system has to be low latency. Because latency is the number one killer. There’s no point of observing something when it’s already done and it’s been 20 seconds. So that’s the key insight that we found as a team that, hey, these LLM systems are pretty nuanced and complex. The outputs of these are very hard to quantify objectively, whether they’re good or bad. So you have to somehow use this technology itself to do evals and observability. And we wrote a whole paper on this, on how you use other LLMs to judge the outputs of LLMs. But the simple paradigm of making another LLM call to evaluate a different LLM simply won’t scale. So that was kind of the aha moment for us to come back to the drawing board and really figure out how do we build low latency systems in this era. And that is one of the things that has become the topmost problem for the Fortune 500 and beyond. Every leader is figuring out how do I productionize my agents today. And I’m not going to fly blind. So no one’s going to put an agentic system, especially if they do actions and tasks, they’re not going to do it without any kind of observability. And if the entire observability playbook fails, that worked for traditional applications, it fails in the new era, what tools do I have to observe these agents? That’s the key gap that we fill. And that is the problem that we’ve been working on, especially focusing on production, because production is a bit of an infrastructure problem. Luna is, of course, one of the key ingredients. And I think there’s two parts to Luna’s innovation. There’s the modeling side, which is how do you tame the intelligence of LLMs into smaller models and use small models efficiently. But a vast majority of the performance gains and latency gains that we see is on the infrastructure side, which are these AI platforms and LLM inference platforms that allow us to host Luna. So for example, we’ve built our own in-house LLM inference engine, which is designed to host small language models. They’re not designed to host the larger models. We’ve done tons of optimizations, including techniques like low-rank adaptation that allows us to run many Luna models under a single foundational base. So you need only a single GPU. But net-net, it allows us to build these evals, which can run at breakthrough latencies, 100 milliseconds and below, which has not been seen by the world. And they are actually capable of evaluating a complex end-to-end interaction between a human and an agent. So that was the key unlock for Galileo.

Siddhartha Ahluwalia 33:15

And if you have to explain, because eval has become very popular in basic terminology, what are evals and why they have become so popular?

Atin Sanyal 33:24

Yes. So eval stands for evaluations. And an evaluation is essentially a metric or a signal that you need to judge whether the output of an LLM is good or bad. That’s the simplest definition. The question is how you do it. Of course, eval spreads into input. There’s all different components to an LLM request, especially if you’re building an application for the app developers out there, they would know that in order to build even a simple application that uses LLMs, there’s some sort of a database lookup, you’ll have to set up a vector store, and you’ll have some system of record. So there’s many different components. And one small issue in any of them can lead to a bad output. So the whole discipline of eval says, how do you prevent bad outputs from going back to the user that can be catastrophic for a business.

Siddhartha Ahluwalia 34:27

Got it. And you have been a big proponent of small language models, right?

Atin Sanyal 34:32

Yes.

Siddhartha Ahluwalia 34:33

Why is that? Whereas, you know, today, the earlier the gap was that LLMs are hallucinating, right? And they can’t be as accurate as small language model. But the way LLMs are getting better every day, they’ll outpace SLMs.

Atin Sanyal 34:52

That’s a very interesting point and an interesting thought. You’re right that LLMs are not only getting smarter, but they’re also getting cheaper to deploy.

Siddhartha Ahluwalia 35:03

Every year, the cost of LLMs is reducing by 80%.

Atin Sanyal 35:08

That’s correct. So they are reducing by 80%. And these models are getting smarter. But there’s a few factors at play here. One is, of course, latency.

Siddhartha Ahluwalia 35:19

And assume latency, they’ll also solve with time. Now, even if you throw complex queries to plot today, the results are usually under a minute.

Atin Sanyal 35:30

No, that’s totally fair. So all these factors are going down. But despite that, there’s some nuances in the economics of this, which you’ll see that it’s despite these improvements, the challenges around evals will still remain. The first reason for that is the more the usage of LLMs in the enterprise and the more gen AI adoption, the infrastructural challenges will keep growing. Because even if the per unit latency reduces, more traffic means you’re just subjecting the system to more tokens and just more data. And all that really brings the system down. That’s one. Number two is agent interactions are getting more and more complex. So even at a unit level, if, say, you apply some magical engineering and solve the problem of, hey, I reduced a 10-second, 15-second complex generation to one second, agentic interactions will compound. Because it’s not about a single request. And the software that we are building, the new era of agentic software, they’re going to get more and more sophisticated, more and more nuanced. That will just keep widening the gap.

Siddhartha Ahluwalia 36:47

Yeah, you’re right. The kind of complexity they’re solving every day is increasing.

Atin Sanyal 36:53

Exactly. The complexity of the use cases are increasing. And really, the thing that LLM providers are optimizing for, the number one thing they’re optimizing for is general reasoning and making the intelligence of these models better.

Siddhartha Ahluwalia 37:13

What do you mean by general reasoning?

Atin Sanyal 37:15

Every new model that comes out, you publish new academic benchmarks and you see that it’s able to achieve a higher score on a particular benchmark. And that just means that it’s able to solve general purpose problems better. And that makes sense for the hyperscalers. And even if they’re getting into enterprises, which we’ve seen a lot of hyperscalers are solving for enterprise use cases, the general techniques that they are, the innovations that they’re doing on the modeling side, there’s a bit of innovators, a little bit of, what do you say, push and pull between, hey, I want to solve a very specific banking use case. But I also want this model to be much better at other things. So you’re going to throw more data into it. You’re going to have more fine-tuned. And yes, you’re going to increase the size of these models. But despite that, there’s many variables which are not in the control of these LLM providers. One is the data. You have no control over the use cases and data is going to change. And it’ll be like, it’s like a fingerprint, right? It’s different for different people. So one of the things that we do with SLMs is not only is it small, so the net economics of it is disproportionately smaller, right? It is so cheap to run a single inference on an SLM versus an LLM. So that gap will always be there. We’re kind of solving the opposite problem of, hey, let the models get better. Let the cost of intelligence reduce over here. It is in our favor, because some of the base models we’re using are getting more in. What we’re doing is we’re adding a layer on top of it, which leverages the specificity of the customer data. And that becomes a very big competitive mode. And then there’s the question of privacy. There’s a lot of hesitation on sending very key proprietary data. It has always been. And even though there’s a lot of cloud adoption, the fear of AI is 10x the fear of putting data on the cloud. People have a lot of trepidations around what is, are they training on my data? And I don’t want to send my data to an AI. It’s such a black box. So that fear is going to be there. Hopefully it will subside over the next maybe 10 years. But even today, 10 years after the whole cloud revolution, there’s still questions around. People still talk about data privacy and there’s all these compliances, etc. So there’s all these different motions at play. We feel like if you really focus on the fundamental problem that, Hey, evaluations for language models is a very task specific constraint problem. And it does not require you to use general purpose reasoning for that. And then there’s fine tuning you can do on the proprietary data. You can achieve really high accuracy with very low resources. I feel like that’s a headwind that will always rather that’s a tailwind that’s always in your favor, no matter how cheaper the models get, there’ll always be better models. But one thing is a fact that there will be more AI usage over the next many years. People are adopting it. Just the net usage of these AI systems is going crazy. We already know OpenAI and Anthropic, their revenues have just catapulted in the last few months and it just goes up.

Siddhartha Ahluwalia 40:51

Anthropic is at $40 billion of run rate, right?

Atin Sanyal 40:53

Right. And they were valued at 18 billion from maybe a year and a half ago. And from there, they’re almost a trillion dollars. And the revenues have kind of seen the same level. So this will keep going up. The technology is very real and everyone realizes it. That’ll increase the traffic. So it’s smart to think about evals and observability in a much more constrained way. And how do you solve it cheaply? Because it’s the cheap and efficient ability of these SLMs that can really help it scale no matter how much the AI footprint.

Siddhartha Ahluwalia 41:32

So you’re saying there’s a huge opportunity building first of all in SLMs for domain specificity and then building evals for those SLMs.

Atin Sanyal 41:44

Well, I’m actually mostly referring to evals and observability, but you’re right. There’s many instances where you can actually train and fine tune an SLM and use that for, say, generational capabilities. But for me, that’s a non-issue in the sense that, hey, you can continue using newer models and you should, because they are genuinely better and they will get better and better at doing more advanced things. So you should not come back to the drawing board and use smaller models for those. If it suits you for your use case, then perhaps, but that’s a decision you should make. But generally, if you’re building any kind of general application system that’s real problems, say CRM or any kind of end-to-end workflow, you should try to disrupt it with these newer models. So my contention is go ahead, use the best models. But observability is a bit of a different ballgame because it has a different set of challenges. It’s an infrastructural problem and latency and cost and quality. Those are the three things. And they are a bit of opposing forces with each other. It’s like this triad of low cost, low latency, high quality. You can only get two of the three if you use traditional techniques. So it’s very important to find that equilibrium where how can you leverage high accuracy of evaluations and observability at low cost and low latency.

Siddhartha Ahluwalia 43:14

Got it. And today, let’s say if your revenue is X, how much of it comes from observability versus evals?

Atin Sanyal 43:23

So it’s kind of the whole online-offline disparity. And this has been there since the dawn of machine learning. And one of the key problems is online-offline parity. That, hey, I built my agent in my experimental environment and now I put it into production and it’s not behaving the way I tested it. You know, it worked on my local host is the famous saying.

Siddhartha Ahluwalia 43:48

The open cloud works on my machine.

Atin Sanyal 43:50

Exactly. It worked on my machine is the last thing they said before it all blew up. So it’s a common problem. It has been there for years. It’s only a bigger problem with agents because they are harder to evaluate. So coming back to your question, there’s evals. Evals is kind of the foundations of it, but offline evaluation is what we call it, which is the whole experimental setup to get to a first version of your agent that you are confident of. And then there’s online observability, which is how do you measure in real time in production?

Siddhartha Ahluwalia 44:28

Either across your enterprise or at your client’s destination.

Atin Sanyal 44:32

Exactly. Exactly. One of our realizations, and we were early to realize this at Galileo, is that online and offline observability are two parts of the same coin. In fact, it’s not a point in time workflow that you do. They are this connected flywheel. And I’ll explain to you what this means. So say you’re an app builder, you build an agent and you test your agent on local host. What you should do, what is good practice is define the behavior of your agent. What all do you want it to do well? And what does good mean? What does bad mean? Define them in the form of evals is what we call it. So an eval is this entity, which is giving you a quantified output of good or bad. So you build a bunch of these evals and let those evals fail. In fact, you build the evals before you build the app. That’s why people say that, hey, evals is the new weapon for product managers, because they are the ones who are defining the app’s behavior. So there’s this notion of evals-driven development in the offline setting where an app builder is building their app and let the evals fail. Let the product managers or the builder themselves set up these evals and let them all be red. And as you build, let them auto turn green. And once they’re all green, you know that you have some quantified sense of, hey, this is a version of the app that I’m confident of. Then you ship it to production. The production will likely meet some of the eval criteria you had, but it will very likely meet newer criterias which you did not test on. In Galileo, we call them unknown unknowns. So your agents will fail in nuanced ways where you didn’t expect it. That’s where the discipline of catching these unknown unknowns and creating new evals and bringing them back to the offline testing environment. That’s what creates the flywheel. And there’s many other elements to this flywheel where you have to collect the data and look at the traces and evaluate them, perhaps manually through a subject matter expert, but you can also use offline judges to evaluate them. And these are workflows that we are seeing already in banks, in telecom companies, in consumer good healthcare. They’re all starting to build these individual workflows, but it’s all disconnected. That is the problem. And that is specifically the problem that we’ve solved at Galileo. How do you build this connective tissue of doing evals, then measuring new evals or unknown unknowns, which are fodder for the new evals. And then you do more testing and then release and then catch issues that are false positives in your evals and have human feedback in the loop. And then as you scale your agent, you know, these evals, which are LLM judges, they will not scale and collect the data as the foundational training data for these SLMs. And then you compare the SLM performance with the LLM judge performance. And then once you see parity, then you productionize the SLM. And now you have full-scale observability at throwaway cost. This is this whole loop. This is Galileo in a nutshell.

Siddhartha Ahluwalia 47:46

And can you name me without sharing the name of the customer, like one industry or one specific customer from our industry where you saw agents in production that are doing autonomous tasks across various verticals or various use cases using Galileo?

Atin Sanyal 48:11

That’s a great question. Certainly, I can say the bad news is that the reality is it’s still very early.

Siddhartha Ahluwalia 48:19

Yeah.

Atin Sanyal 48:19

Even in the enterprise, folks are building simple agents. The enterprise kind of evolved from building sort of knowledge worker workflows back in 2024, 2025.

Siddhartha Ahluwalia 48:31

What is a knowledge worker?

Atin Sanyal 48:32

Like internal tools that make employees more efficient. So having, you know, kind of like a company’s like Glean, right?

Siddhartha Ahluwalia 48:39

Internal search for everyone.

Atin Sanyal 48:41

Internal search. So that’s been kind of the hot use case for most companies.

Siddhartha Ahluwalia 48:44

But that was an old hot use case.

Atin Sanyal 48:45

It’s an old hot use case. So now we’ve evolved into one. How do you make these use cases more agentic? Still internal tools, but they do more things than just Q&A lookups, like creating a graph or creating a presentation. That’s the new sort of internal efficiency workflows. But there are a set of customers who are sort of living on the edge. Of course, you can put beta stamps and stuff in your product, which they do for AI features. But a lot of the… Some of the more cutting edge companies that use Galileo, the use cases I’ve seen are these agents that take… Use tool calls and take action and get data from different APIs and really surface much more nuanced, actionable information to the user. And that’s really baked into the product. Without naming a customer, a specific example, there’s a sales intelligence platform that uses Galileo for observability. And the platform, what it does is it gives you sort of enterprise customer intelligence. It’s mainly for go-to-market teams to discover customers. And they’ve built this beautiful Q&A chatbot-like system where the interface is still language, but it’s doing a lot more things than just answering questions from a database. It’s going into using various APIs and open source CRM tools to get information from… About what employees are doing and what they’re posting on social media, all this very interesting information. And you can auto-create these sales battle cards and save them and bookmark them. And it’s really easy to use and very intuitive. I really love that product. But behind the scenes, they’re running all these autonomous agents. So this is one example. Other examples are like AISRE. That’s another very hot use case on detecting root cause RCA issues for traditional applications. So they use us as sort of their monitoring their agents. The most cutting edge use cases I’ve seen, they use our agent control product.

Siddhartha Ahluwalia 51:00

What’s an agent control product?

Atin Sanyal 51:02

So agent control is a recent product that Galileo launched. And as the name says, it is to control agents or provide a control plane for seeing what the agents are doing at a particular point in time. It is open source. So you can Google agent control Galileo, you’ll find the open source repo. And what this does is it’s meant to design these policies, which you can apply to specific steps with the agent. And those policies will either shut down the agent or make it reroute to a different tool call. And what I’ve seen is it makes the net-net agentic experience a lot more reliable and stable. Of course, there’s that balance you want to find between… It should not be too unstable and stodgy where it says, I don’t know everything. But it’s this control plane that allows you to control these agents. So some of the more cutting edge use cases are these long running agents, which are doing longer workflows like end-to-end. Think of any enterprise workflow, whether it’s a financial workflow or a CRM workflow, they’re attempting to do those with agents end-to-end. And the agent control serves as the control layer for essentially observing what’s happening at each step.

Siddhartha Ahluwalia 52:20

Your customers include Reddit, Airbnb, P&G, Comcast, and six of the Fortune 50. How did you build your GTM? And especially, all of us coming from India are hardcore engineers. So we are not really taught sales back at home.

Atin Sanyal 52:37

Absolutely. So I came to the US in 2011 to do research, was in a research lab at UCLA, did my master’s there. And like you said, I was a hardcore engineer. I loved solving problems, hard problems. So spent 10 plus years in big tech before starting the company. And I feel like I had to do a lot of unlearning of certain habits, which I had built while working in big tech, because not all problems are Apple and Uber level problems, especially in the enterprise. So first six to eight months of our journey was talking to hundreds of people. And this is in an era where we did not have granola and note takers and voice agents. It was all just taking notes in a diary. And we would meet people in coffee shops here in New York, we would fly. So a lot of talking just to really figure out what the problem was. My leverage was that I had built AI systems at scale before and truly had built some of the foundational infrastructure that defined the MLOps industry itself. So I brought that to the table. What I had to unlearn was how do I build something very simple and very narrow that solves a very specific workflow problem. And that is something which I learned over dozens of conversations we had. So most of our first year was spent just doing this. Yes, we would build some prototypes here and there, but it was mostly throwaway. Once we figured out that really, the problem of trust in AI is a big problem, and correlating that with, hey, it is very evident that more and more AI systems in the next three to five years are going to be unstructured AI-based systems, it really gave us the conviction of the specific problem that machine learning and AI is a garbage-in-garbage-out problem. So we have to solve the garbage-in problem.

Siddhartha Ahluwalia 54:50

So for example, let’s say, imagine I am a potential customer, right? Who would my ICP be in your earlier years? It would be the CTO, the CIO, or VP engineering in an enterprise.

Atin Sanyal 55:02

You are the potential customer of Galileo?

Siddhartha Ahluwalia 55:05

No, yes, the potential customer of Galileo.

Atin Sanyal 55:07

So in our case, our ICPs were the VPs of engineering, or rather VPs of data science back in the day. It was a role that apparently is a diminishing breed. But we would go to these advanced AI teams, mostly at larger financial companies and telcos, and we would go to Capital One and talk to data science teams there and really ask them, what are you doing on a day-to-day, and what is the key problem?

Siddhartha Ahluwalia 55:36

Got it. And let’s say initial screening is done, you’re figuring out a problem. Now, I’m potentially like the 20th customer of Galileo. How would your pitch change now?

Atin Sanyal 55:51

That’s a good question. I think the pitch difference depends on the maturity of the product, of course, but also the maturity of the landscape. This area is moving so fast that between customer one and customer 10 and customer 20, just the entire wording of the pitch has changed. But what has not changed about the pitch is this core problem of, hey, I’m building AI systems and agents, and I don’t want to fly blind. How do I measure these AI systems effectively? So give me a system to measure good and bad. So that has not changed. And more than us telling the user, it’s the users who tell us that, hey, I’m looking for a way to quantify good and bad in my data. And this thing has not changed since day zero of our company when we started early 2021. It’s still all about how do you quantify good and bad, even in the era of agents. It’s the same pitch where we say, our pitch is, we are fortunate enough to have a product which almost needs very little selling as far as the problem is concerned, because everyone will tell you that, hey, I’m building AI systems. I need a way to measure things. I’m not going to productionize it. I’m flying blind very clearly. The question is how? I think the challenge is how? But it comes back to us being engineers and deep systems thinkers, which is a great problem. So our main challenge is, hey, how do we build a good product that can solve this measurement problem?

Siddhartha Ahluwalia 57:31

And you recently switched roles from CTO to CPO. What’s been behind that transition?

Atin Sanyal 57:38

Yeah, so I worked on the technology of Galileo. I literally built the first version of Galileo in the earliest days during our pre-seed days. And I’ve always been a builder and engineer and led many teams in engineering, always been in the engineering space, right from the beginning, especially in the AI space. The new era of software, people are figuring out what does it look like with agents, with Gen AI, with tools, and it’s just the beginning. OpenClaw was one blueprint that was published and it just blew up because it said that here’s one step that I’ve taken at how to build a long running agentic system. And it’s only a blueprint. OpenAi. OpenClaw, my apologies, has many flaws. And anyone who’s worked deeply on the system, it is a brilliant idea that’s been put out and the creator of OpenClaw is absolutely fantastic. That said, it can be maneuvered in many ways. And I have seen that personally as I’ve built out on my personal local host laptop. I’ve tried a lot of things and I’ve tried to move things, really break OpenClaw in many ways. And I’ve seen that while it has its flaws and rough corners, it’s the first blueprint of what an AI agentic system could look like. So yes, so I’ve played a lot with newer Gen AI systems and I’ve always been on the cutting edge. The form factor, people are still figuring out of what product in the new age of AI looks like. And there’s some ideas that yes, you need still some sort of a browser form factor. Some of it is automatable. There’s agents in the loop. There’s the CLI form factor, which is hot again. And you have cloud code and other sort of CLI applications. There’s MCP connectors where you have a chatbot. So it’s all sort of taken the world by storm that, hey, traditional enterprise applications, which used to be a dashboard, which would show a chart with squiggly lines on it. All that era is over. Now it’s the era of actions and intelligence. So I took on the product, a hat, I wore my product hat because I’m a deep sort of engineer and engineering leader. I’ve come from this space, I at least understand it from silicon up. So it’s a good role to take in an era where builders are the kings and queens. And me being just having that builder DNA, I think it helped us do more experiments, think about what good product looks like. Because even the observability platform, it is a product. And it needs to be built out really well. We have this internal term we use called umami. And umami means something that just feels so good. And in an era where you can build a website and a web application with a Postgres database behind in 30 seconds, software is just commoditized. Software is just tokens, code is tokens. So how do you build high quality product in the new era was a key problem that we had to solve, take a step back, think about outside of all the machine learning and AI innovations we do, what is the new product? That was the key question. I took on that role to try to answer that question.

Siddhartha Ahluwalia 1:01:15

So when software has become tokens, what does the final definition of a great product be?

Atin Sanyal 1:01:23

That’s a great question. I was figuring that out on my way here. And the day before that, and the week before that, that’s pretty much all I think about. I can certainly share my, I guess, my learning still now. It is, there’s no, no one knows the right answer per se. It is, what I’ve realized is that you have to meet people and workflows where they are. Because it’s not a clean sweep, clean canvas, right? There’s software is pretty hairy. If you really look at what goes on behind the scenes, there’s a lot of traditional software that’s already in place, which solves really hard problems. So we have to fit into that and then evolve the form factor. The easiest thing to disrupt is what a good UI looks like in the new era. And I can certainly tell you about in the observability and eval space, what is the problem that any observability tool really solves, whether it’s AI or not AI? It’s the, hey, what’s wrong question. You ask a system, hey, what’s wrong? And then you take 50 steps to figure out what’s wrong. Traditionally, we have solved that through statistics, through charts, through drift, through all kinds of mathematical paradigms that we put into a software. And that shows up like a chart. And that’s what developers do. Now you’re in the era where you can ask a question and get an answer. And there’s all these steps in the middle, which requires reasoning. You don’t have to build custom business logic for that. And the primary interface for this question, hey, what’s wrong, can be voice and can be language, text, because those are the primary modes of human interaction. So it all starts with that. And then from there, you trigger this workflow where your little agents who are little workers, they do their things and they bring you the results. So what’s in between is still to be figured out. What I know is that the current form factor has to fit into the existing form factors, which is some sort of a browser, some sort of application. There is a case to be made for a natural language-based interface, whether it’s voice or language. Certainly, a lot of voice-based interfaces are also getting really popular. We’ve seen this movement of voice interfaces becoming the standard as the quality of voice models are getting better. You’re talking to one of my friends later on, who is running a pioneering company in voice. And a lot of these businesses are doing really well for this reason that voice quality is so good. So in the end, we’ll enter an era where we’ll talk to software and software will do work for us. And everything in the middle is a means to an end, whether it’s a blue button at the top right, or whether it’s a chat bot or ChatGPT-like interface, it will likely be some sort of a hybrid where a button makes sense. We don’t want to replace it. But ideally, maybe 10 years down the line, or perhaps 5 to 10 years down the line, if we can build a system which is completely headless, where you just start kind of going back to Iron Man, where you just talk to the walls and they answer you and they get stuff for you. That’s the ideal thing. And then taking it one step further, actually fixing issues for you. So we are going towards that trend, but it’s going to be meeting the world where it is and taking, it’s like a ship of thesis, replacing one brick at a time.

Siddhartha Ahluwalia 1:04:57

Got it. And, you know, last of all, big congratulations on the exit, right? A monumental life journey, you know, and milestone in the life of a founder, right? And getting acquired by, let’s say, a Fortune 50 company like Cisco, which has been a leader for the last many decades, right? So how did the marriage with Cisco happen? Can you share the process, whatever is shareable?

Atin Sanyal 1:05:25

Absolutely. So, you know, the work with Cisco that we were doing even before we got acquired by Cisco was super interesting. So we really met them much before today. And the relationship kind of evolved into that it really makes sense for them to acquire us. We started maybe a year and a half ago working on certain interesting sort of future-facing problems on how do you define agent-to-agent interactions and, you know, in the era of then A2A and-

Siddhartha Ahluwalia 1:06:02

That was all by MCP, right?

Atin Sanyal 1:06:04

Well, MCP is one stab at solving a very specific problem of agent interactions, which is selecting the right tools and sort of automating that process. There’s a lot more to be solved. It’s a drop in the ocean of the overall set of interactions. So anyway, there were some open-source efforts, specifically agency that we partnered with them on and did a bunch of other sort of similar things. So we knew the team. Splunk, particularly, it was acquired by Cisco a couple of years ago, and they are an absolute beast. They are kind of running the data systems of the world. They have thousands and thousands of enterprise customers who use them as a telemetry logging system. But also over the years, they’ve built out this amazing observability stack, which is kind of silicon up. They have observability at each layer. Cisco’s traditionally known for networking and security, but they have really honed in on application observability in the last decade. And the gap was agents. So now traditional applications will move. I said, right, software is just tokens now. So everything is just gonna change very quickly. So that kind of shows you the sense of urgency. That’s why we are such a clean fit into Splunk’s overall stack. But not just that, Cisco is one of the fewest companies in the world, which is treating observability and security as two sides of the same coin. They have products such as AI defense, which is sort of more security. They have network security companies, acquisitions that they’ve made, like Thousand Eyes in the past, AppDynamics, of course, Splunk. So they’ve filled out a lot of these gaps that they had, and they’ve built this beautiful stack on top of their federated data settings. So we are sort of their agentic observability and security play. So that’s how it kind of happened. I feel like looking at the five-year journey, going back to the first day, I mean, I can’t even recognize the company that it is because the whole world has changed four times over in the last five, six years. So firstly, thank you for the wishes. It certainly means a lot as a founder to see the company succeed and seeing all the employees, all the work that our engineers, our product teams, our GDM teams, all the work they’ve done come to fruition. But it’s really not the end of the Galileo story. We will continue the Galileo story just within Cisco, except now we have the rocket fuel of unimaginable proportions. So I’m personally super excited. A company which was literally started in my garage with a little table, and there was a heater because the garage was too cold. Going from there to hopefully the next five years, Cisco serves 80% of the world’s internet. And it’s just unfathomable the scale at which they operate. They’re a phenomenal company. They’re the top 25 companies in the world. And to see Galileo be the agentic fabric for any application that’s on the Cisco stack, I mean, it’s a dream come true.

Siddhartha Ahluwalia 1:09:25

Yeah, absolutely. Congratulations again, Atin. Thank you. And thank you so much again for the podcast.

Atin Sanyal 1:09:29

Thank you for having me. This was such a pleasure.

Siddhartha Ahluwalia 1:09:31

I think it was one of the most beautiful conversations that I ever had.

Atin Sanyal 1:09:36

Thank you.

Vector Graphic Vector Graphic

Know when new episodes are released. Subscribe to our newsletter!

Please enter a valid email id