
Top 3 Winners:
Nishant Shinde, Akash Aswar, Aniket Hupele
Congratulations!!
1. Which statement best summarizes MLflow's role for GenAI applications?
• A. It is only a model-training library
• B. It provides capabilities for tracing, evaluation, prompt management, and model/prompt comparison
• C. It only stores source code
• D. It eliminates the need for monitoring
✔ Correct Answer: B. It provides capabilities for tracing, evaluation, prompt management, and model/prompt comparison
2. Which sequence best represents the AI application lifecycle?
• A. Build → Deploy → Code → Test → Delete
• B. Build → Track → Trace → Evaluate → Optimize → Deploy → Monitor
• C. Train → Predict → Store → Delete
• D. Prompt → Model → Database → API
✔ Correct Answer: B. Build → Track → Trace → Evaluate → Optimize → Deploy → Monitor
3. What is the primary reason AI/LLM applications need continuous tracking and tracing compared to traditional software?
• A. They are always open-source
• B. They have probabilistic outputs and behavior can change without code changes
• C. They require more developers
• D. They run only on GPUs
✔ Correct Answer: B. They have probabilistic outputs and behavior can change without code changes
4. What is MLflow primarily used for in the context of GenAI applications?
• A. Only model deployment
• B. Tracing LLM calls, evaluating responses, managing prompts, and comparing versions
• C. Only data preprocessing
• D. Only hyperparameter tuning
✔ Correct Answer: B. Tracing LLM calls, evaluating responses, managing prompts, and comparing versions
5. Which of the following can be included in an MLflow trace for an LLM application?
• A. User request, prompt construction, retrieval calls, tool invocations, model calls, intermediate outputs, final response
• B. Only the final response
• C. Only the model configuration
• D. Only the dataset name
✔ Correct Answer: A. User request, prompt construction, retrieval calls, tool invocations, model calls, intermediate outputs, final response
6. Why is tracing important for debugging AI agents?
• A. It shows which step caused a failure or increased latency
• B. It replaces the need for code
• C. It only shows the cost
• D. It is only for visualization
✔ Correct Answer: A. It shows which step caused a failure or increased latency
7. Which evaluation approach uses predefined facts to check correctness?
• A. Human feedback only
• B. Golden datasets with expected facts
• C. Random sampling
• D. Ignoring outputs
✔ Correct Answer: B. Golden datasets with expected facts
8. Why is it important to measure latency and cost together with quality?
• A. Because high quality with very high latency/cost may not be production-viable
• B. Because latency and cost are irrelevant
• C. Because quality is the only metric
• D. Because it makes reports longer
✔ Correct Answer: A. Because high quality with very high latency/cost may not be production-viable
9. Why should prompts be versioned like code?
• A. Because they are production assets that affect behavior and need reproducibility
• B. Because they are never changed
• C. Because only models matter
• D. Because it is a legal requirement
✔ Correct Answer: A. Because they are production assets that affect behavior and need reproducibility
10. What is a benefit of prompt versioning in MLflow?
• A. It allows comparison of prompt performance across versions
• B. It makes prompts longer
• C. It removes the need for evaluation
• D. It only stores the latest prompt
✔ Correct Answer: A. It allows comparison of prompt performance across versions
11. Which MLflow feature helps manage prompt and model versions centrally?
• A. MLflow Tracking only
• B. MLflow Model Registry and Prompt Management capabilities
• C. MLflow Projects only
• D. MLflow UI themes
✔ Correct Answer: B. MLflow Model Registry and Prompt Management capabilities
12. What is the role of evaluation results in prompt versioning?
• A. They are irrelevant
• B. They provide evidence of how well a prompt version performs on quality, safety, latency, and cost
• C. They only show the prompt text
• D. They are used only for training
✔ Correct Answer: B. They provide evidence of how well a prompt version performs on quality, safety, latency, and cost
13. What does experiment tracking help an AI engineer achieve?
• A. Reproduce and compare experiments
• B. Automatically eliminate all hallucinations
• C. Increase internet bandwidth
• D. Remove the need for evaluation
✔ Correct Answer: A. Reproduce and compare experiments
14. Two prompts produce different results with the same model. What is the best engineering approach?
• A. Pick whichever looks better without measurement
• B. Compare their evaluation results, latency, cost, and other relevant metrics
• C. Always select the newer prompt
• D. Always select the longer prompt
✔ Correct Answer: B. Compare their evaluation results, latency, cost, and other relevant metrics
15. Which statement best describes systematic AI evaluation?
• A. Judging an answer only by whether it "looks good"
• B. Repeatedly measuring AI behavior against defined criteria
• C. Avoiding datasets
• D. Evaluating only latency
✔ Correct Answer: B. Repeatedly measuring AI behavior against defined criteria
For any queries, please write to capdev@harbingergroup.com