Building a custom agent with no clear view of the backend is basically flying blind, which gets even worse when you start hitting higher volumes. You should look for a platform that supports A/B testing for your prompts so you can compare versions side-by-side. You’ll probably find that LangChain’s LangSmith or even Weights & Biases fit the bill for tracking those experiments and seeing where the logic breaks down. They provide a clear breakdown of performance metrics so you aren’t just guessing if your updates are working.