Field noteAugust 26, 2026

Small AI And Small Talk

A week of conversations in Bangalore

After spending a decade at Apple solving geospatial data problems at scale, I recently left the company. The Bangalore trip last week was my first real work trip since I left my job. There were many conversations around AI. A few stood out.

No Right Answer

"I can fine-tune a model and make it deterministic" - someone came up to me while I was sitting at a co-working cafe in Bangalore and said this a few minutes into our conversation on LLMs. This person now had my full attention. I asked a simple question to the guy with the big claim - Why don't Anthropics and OpenAIs of the world do that every night? Why do they freeze their weights and go web searching on every quest? He didn't have an answer. And even though he didn't, I will anyday have this conversation rather than sit alone in a cafe.

Short term, Long term

Most of my other conversations were around agents - in Bangalore, everyone is building agents. "Define an agent in one sentence", I challenged someone who had been working on agents for more than a year. To me, building agents is a short term problem. At Apple, when I was building the foundational framework for agents, the hard part was convincing ourselves on the determinism, repeatability, verifiability, reversibility of the system. It was also aligning the org on following the same spec. We thought we were building an agentic framework, but what we did was expose the inefficiencies on the system. Agents expect all the rules to be explicitly called out. Big organizations are anything but transparent. The long term problem lies in making the agents work in real environments. And making them dependable is the right start.

Old AI, New AI

The field of AI has existed for many decades. Geoff Hinton, the father of Neural nets has been working on this idea since the 1980s. I have worked in this field since 2012. And the same questions that are asked to LLMs by a larger population have always existed. These questions revolve around the "dependability" of an AI/ML system. Can they be trusted (alignment, hallucination)? Can the result be explained (verifiability, causality)?

At a small gathering focussed on Small AI, I saw that question come up again and again. I am glad it did. My take: Machine learning systems are by design probabilistic. Its more about knowing the confidence (the degree of rightness). Generative AI has changed the narrative drastically. Primarily because they were trained on language of the internet. But that doesn't take away the probabilistic nature of them.

Following two papers on fine-tuning small models showcase that pre-training / fine-tuning results are promising but there is a long way to go. Specifically, Synthetic Continued Pretraining articulates this problem quite well, the same base model with retrieval, untrained, beat the fine-tuned version closed-book (60.35 against 56). Same weights, same size, and retrieval still won. Building Domain-Specific Small Language Models via Guided Data Generation reveals that DPO/RLHF gains are marginal compared to pre-training and fine-tuning. Until SLMs showcase significant improvements, they will need to work in conjunction with rule-based and old ML systems.

My take-aways from the trip

  • Old ways of doing machine learning will not cease to exist. If anything, the awareness of ML is causing more people to seek machine learning solutions. The problem still lies in educating them of it's side effects.

  • The claim that a small language model can be fine-tuned for a specific dataset begs more research and analysis.

  • Small talk is great, it can lead to big ideas. If you are a junior engineer just starting out, don't fear being wrong at the cost of being silent.

Next Steps

All in all, it was a wonderful trip. A lot of questions that I had about myself were answered when I was talking to other people (when I introduced myself or explained my work). I will be spending the next few months exploring two questions in my research: first, the fine-tuning question on Language Models and second, the work that lies at intersection of old school ML models and Language models. The end goal is to make AI more dependable on the edge.

This trip has given me a re-assurance on my path forward. Looking forward to the next one.