ProCoder Quiz Sep 2026

MLFlow: Build, Evaluate and Deploy AI with Confidence

Top 3 Winners:

Nishant Shinde, Akash Aswar, Aniket Hupele

Congratulations!!


1. Which statement best summarizes MLflow's role for GenAI applications?

• A. It is only a model-training library

• B. It provides capabilities for tracing, evaluation, prompt management, and model/prompt comparison

• C. It only stores source code

• D. It eliminates the need for monitoring

✔ Correct Answer: B. It provides capabilities for tracing, evaluation, prompt management, and model/prompt comparison


2. Which sequence best represents the AI application lifecycle?

• A. Build → Deploy → Code → Test → Delete

• B. Build → Track → Trace → Evaluate → Optimize → Deploy → Monitor

• C. Train → Predict → Store → Delete

• D. Prompt → Model → Database → API

✔ Correct Answer: B. Build → Track → Trace → Evaluate → Optimize → Deploy → Monitor


3. What is the primary reason AI/LLM applications need continuous tracking and tracing compared to traditional software?

• A. They are always open-source

• B. They have probabilistic outputs and behavior can change without code changes

• C. They require more developers

• D. They run only on GPUs

✔ Correct Answer: B. They have probabilistic outputs and behavior can change without code changes


4. What is MLflow primarily used for in the context of GenAI applications?

• A. Only model deployment

• B. Tracing LLM calls, evaluating responses, managing prompts, and comparing versions

• C. Only data preprocessing

• D. Only hyperparameter tuning

✔ Correct Answer: B. Tracing LLM calls, evaluating responses, managing prompts, and comparing versions


5. Which of the following can be included in an MLflow trace for an LLM application?

• A. User request, prompt construction, retrieval calls, tool invocations, model calls, intermediate outputs, final response

• B. Only the final response

• C. Only the model configuration

• D. Only the dataset name

✔ Correct Answer: A. User request, prompt construction, retrieval calls, tool invocations, model calls, intermediate outputs, final response


6. Why is tracing important for debugging AI agents?

• A. It shows which step caused a failure or increased latency

• B. It replaces the need for code

• C. It only shows the cost

• D. It is only for visualization

✔ Correct Answer: A. It shows which step caused a failure or increased latency


7. Which evaluation approach uses predefined facts to check correctness?

• A. Human feedback only

• B. Golden datasets with expected facts

• C. Random sampling

• D. Ignoring outputs

✔ Correct Answer: B. Golden datasets with expected facts


8. Why is it important to measure latency and cost together with quality?

• A. Because high quality with very high latency/cost may not be production-viable

• B. Because latency and cost are irrelevant

• C. Because quality is the only metric

• D. Because it makes reports longer

✔ Correct Answer: A. Because high quality with very high latency/cost may not be production-viable


9. Why should prompts be versioned like code?

• A. Because they are production assets that affect behavior and need reproducibility

• B. Because they are never changed

• C. Because only models matter

• D. Because it is a legal requirement

✔ Correct Answer: A. Because they are production assets that affect behavior and need reproducibility


10. What is a benefit of prompt versioning in MLflow?

• A. It allows comparison of prompt performance across versions

• B. It makes prompts longer

• C. It removes the need for evaluation

• D. It only stores the latest prompt

✔ Correct Answer: A. It allows comparison of prompt performance across versions


11. Which MLflow feature helps manage prompt and model versions centrally?

• A. MLflow Tracking only

• B. MLflow Model Registry and Prompt Management capabilities

• C. MLflow Projects only

• D. MLflow UI themes

✔ Correct Answer: B. MLflow Model Registry and Prompt Management capabilities


12. What is the role of evaluation results in prompt versioning?

• A. They are irrelevant

• B. They provide evidence of how well a prompt version performs on quality, safety, latency, and cost

• C. They only show the prompt text

• D. They are used only for training

✔ Correct Answer: B. They provide evidence of how well a prompt version performs on quality, safety, latency, and cost


13. What does experiment tracking help an AI engineer achieve?

• A. Reproduce and compare experiments

• B. Automatically eliminate all hallucinations

• C. Increase internet bandwidth

• D. Remove the need for evaluation

✔ Correct Answer: A. Reproduce and compare experiments


14. Two prompts produce different results with the same model. What is the best engineering approach?

• A. Pick whichever looks better without measurement

• B. Compare their evaluation results, latency, cost, and other relevant metrics

• C. Always select the newer prompt

• D. Always select the longer prompt

✔ Correct Answer: B. Compare their evaluation results, latency, cost, and other relevant metrics


15. Which statement best describes systematic AI evaluation?

• A. Judging an answer only by whether it "looks good"

• B. Repeatedly measuring AI behavior against defined criteria

• C. Avoiding datasets

• D. Evaluating only latency

✔ Correct Answer: B. Repeatedly measuring AI behavior against defined criteria




For any queries, please write to capdev@harbingergroup.com

Last modified: Friday, 25 September 2026, 3:21 PM