Skip to content
Tiatra, LLCTiatra, LLC
Tiatra, LLC
Information Technology Solutions for Washington, DC Government Agencies
  • Home
  • About Us
  • Services
    • IT Engineering and Support
    • Software Development
    • Information Assurance and Testing
    • Project and Program Management
  • Clients & Partners
  • Careers
  • News
  • Contact
 
  • Home
  • About Us
  • Services
    • IT Engineering and Support
    • Software Development
    • Information Assurance and Testing
    • Project and Program Management
  • Clients & Partners
  • Careers
  • News
  • Contact

Smaller, smarter, safer: How to build agentic AI on the right foundation

When it comes to building an effective AI stack, context is king and power isn’t everything it’s cracked up to be.

“Smaller, smarter, safer — this is a bet our company has taken in how we deploy AI internally,” said Ricky Thakrar, head of sales and account management at Zoho, provider of a suite of popular cloud-based software solutions for sales, marketing, and finance.

“I’m on the business side, and so decisions made by our CIO and IT folks affect me directly, and my teams’ workflows and processes,” he added.

Speaking to a room of tech leaders at the CIO 100 Leadership Live New York event last week, Thakrar explained that every company wants the speed of AI-generated work wedded to the quality of human work, even though these two are diametrically opposed. No amount of model upgrades or spend will close that gap, so the only way forward is to architect your way out. Thakrar encapsulated this idea in a simple formula:

  • Smaller: Stop deploying maximum firepower on every task. Many tasks don’t need it.
  • Smarter: The system around the model decides more than the model does.
  • Safer: Verify at the point a mistake gets locked in, not just downstream of it.

He noted that organizations that win with AI won’t be those deploying the biggest, most powerful models or the most sophisticated architecture, but the ones that figure out that the model is the easy part and the right architecture is harder. That means understanding the hardest element, and the biggest differentiator, is building a human system that learns and compounds alongside agentic systems.

To get it right, organizations need to prioritize the context layer. The size of frontier models like the GPT series, Claude, and Gemini mostly exist to compensate for missing context, Thakrar explained. Without enough context, models need to be able to reason harder and infer more about what a user actually means because it doesn’t know the user’s account, process, or history. A rich context layer makes it possible for enterprises to run workloads on much smaller, lower-power models.

“The intelligence moves from the model into the architecture around it,” he said.

A steep learning curve

One of Zoho’s earliest AI agents was a churn management agent to help the account management team detect churn in customer subscriptions. So when a subscription became inactive, the agent would collect context from notes, meeting recordings, and Zoho’s data enrichment tool, then create a summary of reasons the account might have churned, and schedule a call.

“What happened was I got this churn agent a couple months later, already embedded in our CRM, and within a week my team no longer trusted that agent,” Thakrar said. “The reason is we forgot to collect one very key point.”

In Zoho’s CRM, when a customer buys a bundle of products, that bundle is represented as a single line item. That means the status of any products the customer may have previously purchased individually changes to inactive as they’re moved to the bundle. That’s not churn, but it was interpreted it that way. Zoho fixed it in the second version of the agent.

Then a new problem arose. Many potential customers first purchase Zoho products as pilots or sandboxes. As those customers move from pilot to live instance, they close down the pilot versions. And again, the CRM would record that as subscriptions going inactive.

“The trust deteriorates again because everyone got excited for version 2,” Thakrar said.

Sometimes, a certain product might not be the best fit for a customer and Thakrar’s team will suggest the customer move to another product. That’s deliberate churn, not a churn risk.

“You may have a similar story like this where the agent sounds so good, it’s going to do something quick and add value, but it’s missing context from the account managers, and there are so many more pieces we’re still building out,” Thakrar said. “It’s been almost a year and the problem I have is my team still doesn’t trust it. They’ll see [a message from the agent] and go out and do all the research anyway to make sure it gave the correct answer.”

The team is more on top of potential churn, though, but the promised productivity gains have yet to materialize because the agent has to earn back lost trust due to a lack of context.

“My goal for this year is having an AI-assisted customer journey from sales to account management where the handoff is clean, the context flows, and every piece of information we gather about a customer is weighed, identified, and coached so the sales team can close more deals,” he said.

Zoho’s early experience with agents has led to the idea that constrained, context-rich, deterministic architectures consistently outperform expensive models bolted onto fragmented systems. It all comes down to three pillars: routing, harness, and specialization.

Routing

Routing is about sending workloads to the proper model for the job, which entails providing enough context to a given task that a small, cheap model can handle it without the need for spare reasoning capacity to fill gaps.

Frontier models are expensive and companies can burn through a year’s budget worth of tokens in months. But most tasks can be handled by much smaller, more constrained models at a fraction of the cost.

“You don’t always have to pay the frontier guys for every task,” he said. “We’ve observed with some clients that we could save them 95% with a 3 billion parameter model.”

Harness

An AI agent harness is the software infrastructure scaffolding around an LLM that differentiates an agent from a chatbot. It’s what enables an agent to act on tasks rather than simply respond to prompts. A model reasons through a problem and decides what to do about it. The harness connects the model to the tools, systems, memory, guardrails, and execution environments required to perform the actions determined by the model. The term is frequently used more or less interchangeably with orchestration layer.

“It’s the process around the model, which matters way more than the model itself,” Thakrar said.

In benchmark tests, a superior harness on a less powerful model produces better results than an inferior harness on a much bigger model.

For the best results, Thakrar said, it’s essential to understand the deterministic and non-deterministic elements of a given workload, and build that into the architecture. Machines can read, organize, and validate, and they excel at deterministic tasks. Humans, on the other hand, are exceptional at non-deterministic tasks like judging, synthesizing, and deciding.

Those non-deterministic tasks in a process are the ideal point for AI agents to incorporate a human in the loop, what Thakrar calls human harness. He pointed to a stakeholder mapping agent Zoho built for sales as an example, which takes the context of an initial meeting and third-party enriched data like a LinkedIn profile, weighs probabilities, and makes an educated guess about the stakeholder map.

“The initial goal was just to eliminate that task completely from the human workflow,” he said. “The stakeholder map is done, it’s in the folder, and you can look at it.”

But the agent would struggle to capture nuance. The meanings of titles in organizations always vary, and the politics and dynamics of any given meeting can be difficult for an AI agent to discern. Rather than keep feeding the agent data to try to make it intelligent enough to make those determinations, it was simpler and more efficient for the agent to create a proposed stakeholder map and hand it over to a human who could make changes and explain why those changes were necessary.

Ultimately, Thakrar said the agent still saved human team members time because the stakeholder map was usually pretty close, and the corrections also helped the model grow smarter by adding richer context.

Specialization

Specialization is transitioning a process from testing on a frontier model to production on a much narrower, smaller model. Once you’ve proven that an agent can do a job well, you want to stop paying master-craftsman rates to keep doing that one job well.

Specialization is all about capturing your subject matter experts’ best judgement and pattern recognition to build an open-weight, open source, trained, and fine-tuned model that can be deployed in your own data center.

“The true enterprise bet is to keep that orchestration layer, which is your IP and knowledge, in house,” Thakrar said. “You don’t want to host that on someone else’s model. The goal of everyone in enterprise should be to run, train, and host their own models.”


Read More from This Article: Smaller, smarter, safer: How to build agentic AI on the right foundation
Source: News

Category: NewsJuly 23, 2026
Tags: art

Post navigation

PreviousPrevious post:Stop asking AI nicely: Here’s how to get work-ready results every timeNextNext post:AI success requires a full-stack CIO

Related posts

The new value architecture of the AI-native SaaS era
July 23, 2026
Stop asking AI nicely: Here’s how to get work-ready results every time
July 23, 2026
Principles every enterprise must test before the attack arrives
July 23, 2026
AI success requires a full-stack CIO
July 23, 2026
Sovereign AI has become the public-sector CIO’s control problem
July 23, 2026
Monday.com cuts 20% of its workforce to restructure for the AI era
July 23, 2026
Recent Posts
  • The new value architecture of the AI-native SaaS era
  • Stop asking AI nicely: Here’s how to get work-ready results every time
  • Principles every enterprise must test before the attack arrives
  • Smaller, smarter, safer: How to build agentic AI on the right foundation
  • AI success requires a full-stack CIO
Recent Comments
    Archives
    • July 2026
    • June 2026
    • May 2026
    • April 2026
    • March 2026
    • February 2026
    • January 2026
    • December 2025
    • November 2025
    • October 2025
    • September 2025
    • August 2025
    • July 2025
    • June 2025
    • May 2025
    • April 2025
    • March 2025
    • February 2025
    • January 2025
    • December 2024
    • November 2024
    • October 2024
    • September 2024
    • August 2024
    • July 2024
    • June 2024
    • May 2024
    • April 2024
    • March 2024
    • February 2024
    • January 2024
    • December 2023
    • November 2023
    • October 2023
    • September 2023
    • August 2023
    • July 2023
    • June 2023
    • May 2023
    • April 2023
    • March 2023
    • February 2023
    • January 2023
    • December 2022
    • November 2022
    • October 2022
    • September 2022
    • August 2022
    • July 2022
    • June 2022
    • May 2022
    • April 2022
    • March 2022
    • February 2022
    • January 2022
    • December 2021
    • November 2021
    • October 2021
    • September 2021
    • August 2021
    • July 2021
    • June 2021
    • May 2021
    • April 2021
    • March 2021
    • February 2021
    • January 2021
    • December 2020
    • November 2020
    • October 2020
    • September 2020
    • August 2020
    • July 2020
    • June 2020
    • May 2020
    • April 2020
    • January 2020
    • December 2019
    • November 2019
    • October 2019
    • September 2019
    • August 2019
    • July 2019
    • June 2019
    • May 2019
    • April 2019
    • March 2019
    • February 2019
    • January 2019
    • December 2018
    • November 2018
    • October 2018
    • September 2018
    • August 2018
    • July 2018
    • June 2018
    • May 2018
    • April 2018
    • March 2018
    • February 2018
    • January 2018
    • December 2017
    • November 2017
    • October 2017
    • September 2017
    • August 2017
    • July 2017
    • June 2017
    • May 2017
    • April 2017
    • March 2017
    • February 2017
    • January 2017
    Categories
    • News
    Meta
    • Log in
    • Entries feed
    • Comments feed
    • WordPress.org
    Tiatra LLC.

    Tiatra, LLC, based in the Washington, DC metropolitan area, proudly serves federal government agencies, organizations that work with the government and other commercial businesses and organizations. Tiatra specializes in a broad range of information technology (IT) development and management services incorporating solid engineering, attention to client needs, and meeting or exceeding any security parameters required. Our small yet innovative company is structured with a full complement of the necessary technical experts, working with hands-on management, to provide a high level of service and competitive pricing for your systems and engineering requirements.

    Find us on:

    FacebookTwitterLinkedin

    Submitclear

    Tiatra, LLC
    Copyright 2016. All rights reserved.