Assure AI Engineer
Gemini – AssureAI Engineer
2 positions
Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python
About the Role
You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.
What You'll Own
Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.
Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).
Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.
Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.
Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.
Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.
What We're Looking For
5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).
Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.
Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.
Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.
Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).
Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.
Nice to Have
Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.
Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).
Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.
What Success Looks Like
By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.
By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.
Gemini – AssureAI Engineer
2 positions
Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python
About the Role
You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.
What You'll Own
Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.
Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).
Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.
Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.
Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.
Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.
What We're Looking For
5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).
Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.
Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.
Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.
Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).
Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.
Nice to Have
Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.
Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).
Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.
What Success Looks Like
By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.
By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.
Gemini – AssureAI Engineer
2 positions
Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python
About the Role
You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.
What You'll Own
Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.
Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).
Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.
Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.
Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.
Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.
What We're Looking For
5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).
Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.
Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.
Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.
Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).
Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.
Nice to Have
Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.
Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).
Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.
What Success Looks Like
By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.
By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.