Interview question asked to Data Scientists interviewing at GitHub, Hootsuite, Mimecast and other companies. Original question asked: How do you architect a fault-tolerant pipeline that automatically moves data from an unstructured lake into a structured BI semantic layer?.