A multi-stage AI evaluation experiment designed to test how effectively Generative AI can build, validate, analyze, and make recommendations from business financial data. Using a realistic budget-versus-actual environment, I leverage my 30+ years of experience in finance, budget planning, operations, and human resources to independently validate AI-generated financials, identify errors, evaluate financial reasoning, and assess the quality of executive recommendations.
The project combines Excel, financial analysis, business strategy, critical thinking, prompt engineering, and rubric-based AI evaluation to examine both the capabilities and limitations of AI in financial decision-making—and demonstrate where experienced business judgment remains essential.
I used a realistic business scenario involving a privately held advertising company with 500 employees that was moderately profitable but facing some financial headwinds. I asked the first AI model to build the annual budget and actual financial results, which I then reviewed in Excel to check the calculations and identify any errors. After correcting the financials, I provided the validated results to a second AI model and asked it to independently identify the five most significant financial or business issues and provide recommendations.
I evaluated the AI-generated analysis using my 30+ years of experience in Finance, Budget Planning, Operations, and Human Resources, along with my Excel skills, to assess the financial accuracy and overall quality of the recommendations.
Each of the five issues and recommendations was evaluated using a structured 1–5 scoring rubric across five criteria: