Interview question asked to Data Scientists interviewing at Zenefits, Dell, Patreon and other companies. Original question asked: How do you implement shadow deployments or A/B testing to compare the performance of multiple LLM versions safely?.