QA Engineer II
Current
Masterworks — Nabeh AI Division · Feb 2026 – Present
- Created 120+ benchmark ("golden") test cases for evaluating AI agent accuracy, grounding, and hallucination.
- Evaluated precision, recall, and hallucination-related metrics using Promptfoo and DeepEval across multiple insurance AI agents.
- Validated conversational AI / voice-bot functionality including intent recognition, NLP, and speech-to-text.
- Built and maintained an AI test-generation agent for creating test plans, scenarios, test cases, and pass/fail dashboards.
- Extended the agent to execute against applications and generate structured QA/defect reports.
- Developed automated API tests using Postman and Rest Assured across Agile sprints and documented QA findings for stakeholders.
