Starting with old tickets is useful for a baseline, but you'll need to simulate the ambiguity you actually face. A subset of past tickets only tells you if the AI can solve problems you've already documented and closed.
I'd recommend a silent, parallel phase where the AI drafts replies for every new incoming ticket for a week, but a human agent still sends the final response. Compare the drafts to what the agent sent. The critical metrics are the correction rate and, more importantly, the "time-to-correctness" if the AI draft was wrong. A slightly incorrect AI suggestion can waste far more time than a completely wrong one that gets quickly escalated.
For infrastructure, pay close attention to tickets with incomplete data. Does the AI ask for the missing mount point or server name, or does it confidently guess? That's your biggest risk area.