Leveraging my 30+ years of experience in Finance, Human Resources, Analytics, and structured decision-making, I developed a rubric-based framework for evaluating Generative AI outputs across Business, HR, and Financial analysis scenarios. The project demonstrates my ability to develop evaluation criteria, assess AI-generated responses, apply consistent scoring methods, identify strengths and weaknesses, and determine which outputs provide the most accurate, relevant, and actionable insights.
The prompts were designed to generate three independent AI responses for each of three business scenarios: Business Strategy, Human Resources Strategy, and Financial Investment Strategy. Each prompt instructed the AI to use a different approach, reasoning process, and set of recommendations while avoiding unnecessary repetition in structure, examples, or recommendations. For each scenario, I requested one strategic and comprehensive response, one practical and implementation-focused response, and one risk-focused and critical response. The three responses were then evaluated using a structured rubric designed to assess the quality, relevance, completeness, and practical value of each AI-generated response.
The evaluation rubrics were developed based on the specific requirements of each prompt and the criteria needed to assess a high-quality response. Responses were scored against weighted criteria using a consistent rating scale. The scores were then used to compare the responses, identify strengths and weaknesses, and determine which response provided the strongest overall analysis. I also reviewed the responses for unsupported assumptions, missing information, and the ability to distinguish between facts, assumptions, and recommendations.
| Prompt | You are a business strategy consultant advising the CEO of a 500-employee U.S. professional services company. The company expected a 20% productivity gain from its Generative AI investment but has achieved only 5%. Adoption varies across departments, and some employees spend significant time reviewing and correcting AI-generated work.
Analyze the likely causes of the productivity gap. Recommend a structured approach to determine whether the issue involves technology, adoption, training, workflows, management, measurement, or other factors. Propose five actions for the next 12 months and identify key metrics for measuring business value.
Clearly distinguish between facts, assumptions, and recommendations. | | --- | --- | | Rubric | Rubric Criterion: Instruction Following (15%), Accuracy (20%), Relevance (10%), Completeness (10%), Reasoning Quality (15%), Business Applicability (10%), Clarity (10%), Risk Awareness (10%), scored on a 5 point scale, with Excellent (5), Strong (4), Adequate (3), Weak (2) and Poor (1) | | AI Responses | See Below | | Overall Evaluation | Response A performed best overall because it demonstrate more relevance and business applicability, and more complete analysis of the business problem. Reponse B was had no risk assessment, and did not follow the prompt as desired. Response C focused mostly on risk assessment rather than analysis of the likely causes of the business problem. | | Overall Winner | Response A | | Why it Performed Best | - Adherence to the prompt
Final Scorecard
