Overview

Ai Quality Assurance Tester Jobs in Bucharest Metropolitan Area at ALTEN

Title: Ai Quality Assurance Tester

Company: ALTEN

Location: Bucharest Metropolitan Area

ALTEN Romania is part of the ALTEN Group, Leader in IT and Engineering Consulting. We develop innovative and durable technical solutions that fulfill the needs of our local and international partners.

We are looking for a detail-oriented and technically curious QA Engineer to own quality assurance for our AI-powered chatbot and Agent Framework.

You will be at the intersection of traditional software testing and the emerging discipline of LLM/AI evaluation—ensuring our product delivers accurate, reliable, and high-quality responses both after feature releases and following regular data ingestion cycles.

This role starts with structured manual testing and grows into building an automated QA pipeline for continuous quality assurance.

Main responsibilities:

  • Design, write, and execute test cases and test plans for new chatbot features and agent behaviors.
  • Perform exploratory testing to uncover edge cases, unexpected behaviors, and failure modes.
  • Test conversation flows end-to-end across a range of user intents, personas, and input variations.
  • Document bugs and regressions clearly, with reproducible steps and severity classification.
  • Collaborate closely with AI/ML engineers and product managers to validate acceptance criteria.
  • After any enhancements or upgrades, run sanity testing to ensure core functionality is not impacted.
  • Execute regression tests after each data ingestion cycle to detect answer quality degradation, hallucinations, or factual drift.
  • Track and report quality metrics over time (e.g., accuracy, relevance, groundedness, tone consistency).
  • Work with Data and ML teams to triage quality issues—distinguishing between data problems, retrieval issues, and model behavior.
  • Data Quality
  • Define, maintain, and continuously refine a golden test set of representative queries and expected answers used to benchmark chatbot quality—ensuring questions and acceptance criteria are reviewed and signed off by key stakeholders (e.g., Product, Business, domain experts).
  • Test Automation (Mid- to Long-Term)
  • Build and maintain an automated evaluation pipeline that runs on each release and data update.
  • Automate the majority of manual testing tasks.
  • Implement LLM-as-a-judge and/or rule-based evaluators to score chatbot responses at scale.
  • Integrate automated tests into the CI/CD pipeline.
  • Develop tooling for regression tracking, test result dashboards, and alerts on quality drops.
Upload your CV/resume or any other relevant file. Max. file size: 800 MB.