Independent AI Evaluation & Analytics
We help organisations understand whether their AI works, where it fails, what it costs, and how to make it better.
We don't sell AI platforms. We measure whether they're working — independently, rigorously, and without a product to push.
Sound familiar?
"We've built a RAG application. How do we know it's any good?"
→ AI Evaluation
"We've deployed an AI assistant for customer queries. How do we know if it's actually working?"
→ AI Evaluation & Testing
"Our AWS bill keeps climbing. Where is the money actually going?"
→ Cloud Optimisation
"We have masses of customer interaction data and can't see what's happening."
→ Analytics
What we do
01
Independent evaluation of chatbots, RAG systems and agentic AI. Test sets, failure analysis, LLM-as-judge frameworks and model comparison.
02
Turn complex operational and conversational data into evidence people can actually make decisions from. From ETL pipeline design and data cleansing through to dashboards and visualisations — the full journey from raw data to actionable insight.
03
Measure usage, cost, output quality and business outcomes across your AI tooling — from APIs and agents to enterprise AI platforms.
Developing capability04
Identify where cloud spend isn't delivering value. Analyse usage and billing, quantify savings, assess effort and risk, produce an actionable plan.
How we work
Start with the problem you're trying to solve and the evidence you need — not a generic evaluation template.
Agree scope, success criteria, test data, evaluation methods and deliverables before any work begins.
Test the system, analyse failures and investigate where and why it behaves unexpectedly — including edge cases and failure modes that simple benchmarks miss.
Give you the evidence, findings and prioritised recommendations you need to decide what happens next.
Selected work
Developed a medical triage chatbot for the Nurse on Call service using Cognigy, and implemented an automated conversational testing platform enabling flow testing and regression testing at scale.
Worked as a data scientist tuning conversational language and intent models for enterprise banking chatbots used by major financial institutions. Focus on intent classification, model performance and language coverage.
Played a key role in shaping an AI testing and validation platform for LLM-powered chatbots, RAG pipelines and agentic AI systems. Contributions include multi-judge evaluation frameworks, LLM-as-judge approaches and automated test generation.
Analysed an enterprise AWS environment to identify and quantify cost optimisation opportunities across compute, storage, databases and service commitments. Produced a prioritised plan based on potential savings, implementation effort and risk.
About
Founder & Consultant
Phil has worked in technology, analytics and customer systems for more than 30 years. His current focus is AI evaluation, conversational AI and analytics — helping organisations understand whether their AI systems are actually working, what they cost, and where they fail.
He's particularly interested in the gap between what technology is supposed to do and what happens when real people actually use it.
That question has shaped his work across IVR and speech recognition, enterprise conversational AI, data science and now LLM evaluation.
Tromp IT also draws on Belinda Robin's expertise in communications and media, available for engagements where strategic communications is part of the brief.
Ask Phil
Ask our AI assistant — built and evaluated using the same methods we apply for clients.
Hi — I'm Phil's AI assistant. Tell me about your situation and I'll let you know whether Tromp IT could help.