MarTech360 Interview with Cobus Greyling, Chief Evangelist at Kore.ai


"The first step is to be model-agnostic from the start. I always say the harness is owned and the model is rented. Models are deprecated and discontinued on a regular basis, so organisations are almost forced to change models. That is why they should focus on the harness and on which harness they want to use."
Enterprise AI is moving from individual use cases toward agentic systems that can reason, make decisions and take action. What should enterprises rethink in their AI strategy as they make this transition from generative AI to agentic AI?
I was reading a study recently from 1990. The study was discussing the notion at that time that computers were not being successfully implemented in enterprises. The argument was that the return on investment is not measurable and that the increase in productivity does not exist, at least not as predicted.
The study went back and looked at the advent of electricity as a parallel to the advent of the computer. In short, the conclusion was that a lot of hard complementary work is required to usher in a new technology, even if that technology is very powerful. A technology being powerful does not necessarily translate directly into easy adoption.
So if you ask me what enterprises should rethink, I think it is the notion that the technology alone will not do the work. The complementary hard work will still have to be implemented. There are so many parallels between the advent of the computer and AI, hence the study was so fascinating.
A second point I can add is the data element. There are many things that can be taken off the shelf in terms of models, AI harnesses and agentic systems. But the data is the hard-to-reach portion. Enterprise data is messy, distributed and tribal, to say the least. The hard part will be to transform enterprise data into a durable AI data plane that agents can actually reach.
I always say, the model is rented, the harness is owned. So a strategy that still starts with “which model” is already looking at the wrong layer.
As AI agents become more autonomous, governance can no longer rely only on human oversight. What governance capabilities should enterprises build directly into their AI architecture to ensure agents operate within defined boundaries?
Human oversight still matters, but one can argue it is too slow to be the main control once agents can plan, call tools, keep memory and delegate.
But I also want to look at governance from the lens of chatbots and what we're used to... in a chatbot stack the conversation is the only real object. Tools, models, memory, and other systems show up only when the model happens to reach for them, then vanish back into the transcript. That is enough when the system answers a question and waits. Agentic systems are a step up because they do not stop at language. They plan, call tools, keep state, and hand work to other agents, which means those moving parts now change records, spend money, and cross system boundaries.
If they remain hidden inside a chat log, you cannot say what exists, who owns it, what it may touch, or what it did. Treating agents, tools, models, MCP servers, memory stores, and even agent-to-agent calls as first-class assets is the architectural response to that step up: each one becomes a named object with identity, policy, and a trail of its own, so governance attaches to the action, not to the conversation that preceded it.
Also Read: MarTech360 Interview with Tim Rodgers, CEO & Founder at Ace Workflow
From a customer experience perspective, many organizations are still measuring AI success through metrics such as containment, response time or cost reduction. What metrics should enterprises prioritize to determine whether AI is actually improving customer outcomes?
I think customers experience a company as one continuous relationship. They do not think in onboarding, retention, billing, or support. They think in a single thread: I joined, something changed, I need help, I expect you to already know me. The organisation still sees that thread as a sequence of departments and funnel stages, so the customer is handed off, repeated, and re-identified at every boundary. AI changes the feasible design. Agents can carry context across those boundaries, act on the same memory and the same identity, and complete work that used to stop at the edge of a team. That is the opening to rebuild around the customer’s actual journey rather than around the org chart, so the operating model follows the experience instead of forcing the experience to follow the operating model.
Enterprises often struggle with fragmented AI initiatives across customer service, marketing, HR, IT and operations. How can organizations move from isolated AI deployments toward a more unified enterprise AI strategy?
A large enterprise will never be tidy. Customer service, marketing, HR, IT, and operations will always run initiatives at different stages of maturity, on different vendors, with different owners. Unity is not a single platform that replaces all of that. It is a shared way for those systems to discover each other, pass work, and stay governable. That is where AI is different from the last generation of integration programmes. Protocols such as MCP and A2A, together with a common identity, tool registry, and policy layer, make operability the strategy: an agent in service can call a capability in operations without a year-long programme, and a marketing agent can use an approved enterprise tool without being rebuilt. The centre should standardise the contracts, the catalog, and the control plane. The edges can keep moving at their own pace. Fragmentation remains. Isolation does not have to.
With AI agents increasingly interacting with customers and employees directly, how should organizations think about maintaining trust, consistency and brand experience across AI-powered interactions?
I would say that things like trust, consistency, and brand are properties of what customers and employees actually experience, turn after turn, across agents and channels.
The organisation should treat every transcript and interaction as the control signal: what was promised, what tone was used, where the agent improvised, where it contradicted policy, and where the human had to repair the moment.
So that loop is how brand becomes operational. The feedback from real conversations is what keeps it true: sample the traces, score them for resolution, effort, accuracy and voice, then push the findings back into the catalog, the tools, the memory rules, and the evals. An agent that cannot learn from its own interactions will drift and an enterprise that reads those interactions as a product signal can keep the experience recognisably itself even as the agents multiply.
Looking ahead over the next two to three years, what do you believe will separate enterprises that successfully scale AI from those that remain stuck in experimentation and pilots?
If I can try and put down four key considerations...the first being, treat AI as an operating model, not a feature. Successful organisations will build a durable control plane (routing, policy, evals, identity, tool access, etc ) and a data plane (retrieval, context, feedback, lineage), while pilots stay stuck when every use case is a one-off demo with no shared runtime.
Secondly, own the data and context loop. Scaling fails when models are rented but enterprise judgment stays trapped in messy docs, tickets, and silos. Separators invest in governed retrieval, labeling, memory, and production traces that turn usage back into better systems.
Instrument for “works in production,” continuous evaluation, release gates, drift monitoring, and cost/latency budgets become (or should become) as normal as CI/CD. Experimenters optimize prompts; scalers optimize harnesses and measure quality at parity under real load.
Start by putting agents inside workflows with hard constraints. Successful enterprises run multi-step agents where policy, tools, and handoffs are enforceable not free-form chat bolted onto SaaS. Experimentation dies when the LLM can override the business rules. As Andrej Karpahty says, agency lives on a spectrum and depending on the task, agency should be scaled up or down. He calls it an agency slider, and I like the idea of bounded autnomy for intial projects.
Why chasing the model leaderboard can actually slow enterprise AI adoption by keeping teams perpetually in evaluation mode instead of getting applications into production.
A few things are happening on the model front at once. Multi-model orchestration is becoming commonplace: use the best model for the intended step and use case. There is also a growing pattern of using open models for production tasks and frontier models for development and research. Two recent open models, Muse Glimmer and Nemotron 3.5 Lightning, are focused on locally long-running agentic tasks and are being branded as always-on models.
The practical problem with leaderboard chasing is that the board keeps moving. Teams stay in evaluation mode, swapping models instead of shipping applications. In an enterprise setting, the win is not the next benchmark score. It is a working system in production, with the right model in the right place in the workflow.
How enterprises can judge a new model against their own workflows, requirements and economics, rather than relying on benchmark performance.
Two things matter here. There have been a number of good research papers and a benchmark called Harness-Bench. The gist is that models behave very differently depending on the harness they run in. We need to move from token maximising to harness maximising. That research showed that with the right harness for a specific model, performance can almost double and cost can be halved. Harness-Bench also showed that certain models perform better with certain harnesses. So models now need to be tested against the environment and the task they will run in, and against the harness that will be available to them.
It has become common knowledge that Anthropic, OpenAI, and most likely the other frontier model providers as well, use extensive harnesses behind the API that the user never sees. Often what looks like a model improvement is actually a harness improvement behind the scenes.
It is also important for organisations to run their own AI evals. They already have what they need: user personas, use cases, and different usage scenarios. It is quite straightforward for them to run evaluations and benchmark where models fit into their current environment. This should be commonplace for every enterprise.
How to build a model-agnostic AI architecture that allows organizations to take advantage of better models without rebuilding applications from scratch.
The first step is to be model-agnostic from the start. I always say the harness is owned and the model is rented. Models are deprecated and discontinued on a regular basis, so organisations are almost forced to change models. That is why they should focus on the harness and on which harness they want to use.
It makes sense to start with a simple harness and build it out from there, or to build a harness around an existing model and then run evals to see how that harness holds up against other models. A harness includes the incoming process, memory, context, tools, integration, state management, and so on.
When it actually makes sense to switch models
Look at the methodology OpenAI uses behind its Deep Research API. Smaller language models handle disambiguation. Another small language model clarifies the prompt. A fair amount of work happens before the user request ever reaches the more expensive research model.
An organisation should do the same: use different models for different steps and different use cases. That comes back to AI evaluations. Enterprises already have the use cases, the user personas, and the different scenarios. If they also have conversation and session logs, they can use those too.
We will see AI evals together with harness engineering become much more prominent. Switch a model when your own evals, economics, and workflow fit show a clear gain, not because a public leaderboard moved.
Thanks Cobus!
Cobus Greyling is an AI and conversational technology professional who focuses on the intersection of artificial intelligence and language. As Chief Evangelist at Kore.ai, he explores and communicates developments across large language models, AI agents, agentic applications, conversational AI, and related development frameworks. His professional work includes writing, industry discussions, and educational sessions focused on how emerging AI technologies can be applied to enterprise use cases. His background also includes product-focused roles in conversational AI and voice technology, giving him experience across both the technical and practical aspects of AI adoption.

