Search Jobs

Search by job, company or skills

Senior AI Quality Engineering Innovation Engineer

Senior AI Quality Engineering Innovation Engineer

cloudmarc
5-7 Years
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Senior AI Quality Engineering Innovation Engineer

Offshore, Philippines. Building and deploying AI agents, agentic workflows and automation for a major Australian enterprise testing function.

Employer

CloudMarc Pty Ltd, Melbourne, Australia

Engagement

Permanent or long-term contract, start as soon as possible

Hours

Australian Eastern business hours, overlapping Melbourne from at least 09:00 to 15:00 AEST

Location

Philippines, remote, full time and exclusive to CloudMarc

Reports to

Technology Innovation Director and Head of AI, CloudMarc

The role

CloudMarc is a Melbourne consultancy specialising in quality engineering and applied AI for large Australian enterprises. We hold a funded backlog of innovation work for one of our largest clients: AI agents, agentic workflows, automation and data that change how testing and quality engineering are done across a large, integration-heavy technology estate. You will build those agents and workflows, prove they work, measure them, and turn the ones that pay off into repeatable patterns the client's teams adopt.

You will work from a prioritised backlog with CloudMarc's Head of AI as your single point of contact for the client, and you will own each item end to end: understand the question, build the mechanism, measure the result, write it up so someone else can apply it, and hand it over. Nobody will hand you a step-by-step plan. You will be expected to work out what needs doing, learn what you do not yet know, form a view, and make it happen.

Senior individual contributor. Not a test management, consulting or people-management seat, and not client-facing.

The kind of work

Each item is three to six days of continuous effort with a stated outcome and a measure. Most of the backlog is agentic; the rest is the engineering that makes agents usable in an enterprise.

  • Agents and agentic workflows on an enterprise-approved platform. Designing, building, evaluating and deploying agents that read requirements, code, configuration, test assets or documentation and produce something a testing team acts on: requirements-quality findings, documentation of legacy systems, interface and contract inventories, test plan and report drafts, script reviews on merge requests, knowledge-transfer assistants. Structuring them as reusable skills with knowledge and steering layers, memory that persists across sessions, a human check before output is used, and an evaluation that says how often they are right. Designing each one so a later orchestrating agent can call it.
  • AI-assisted quality engineering practice. Bringing agents into the daily work of test teams: authoring, failure triage, selective regression, review of AI-generated tests, and the habits that make a team trust the output.
  • Pipeline and platform work. Approval steps, notifications, deployment registers, on-demand execution and automated review stages in CI/CD pipelines; the plumbing agents need to run unattended.
  • Quality data and metrics. Extracts from work-tracking and test-management tools through their APIs, measures such as defect leakage, panels that need no manual entry.
  • Test automation foundations. Starter frameworks, API and integration coverage, regression packs, test-data tooling, groundwork for contract testing.

What you need to bring

We are not looking for every skill on the backlog. We are looking for someone who has already built real agentic capability, has an engineering foundation underneath it, and has the agency and curiosity to close the rest quickly.

  • Agentic engineering, with evidence. You have designed and shipped agents or agentic workflows that people other than you use: a trigger, a knowledge base or retrieval layer, skills or instruction files, structured output, a review gate, and some measure of how often the output is right. You can walk through one real run end to end and say what it got wrong and what you changed. You know the difference between a prompt and a system, and between a demo and something a team trusts. Daily use of GitHub Copilot, Claude Code, Codex or an equivalent is assumed. The agent work here runs on GitHub Copilot in an enterprise setting today (agent mode, custom instructions, repository-based knowledge) and the client's platform will change over time, so what matters is platform-agnostic design: the durable parts of a workflow (skills, knowledge, evaluation set) built so they move from one vendor to the next. Experience on any comparable platform transfers; no specific vendor is required.
  • Quality engineering with code. Five or more years hands-on. You write TypeScript or Python (and can read Java or C#), and you have built or substantially shaped an automation framework (Playwright preferred) rather than only written scripts inside one.
  • API and integration testing, and an understanding of microservices. You test REST and messaging interfaces directly, you understand how services, queues, contracts and middleware fit together in an integrated estate, and you can reason about where a defect between two systems should have been caught.
  • Practical CI/CD, GitLab preferred. Pipelines written in GitLab CI or a close equivalent (GitHub Actions, Jenkins, Azure DevOps), with tests, approvals or notifications wired in. Productive in GitLab within your first fortnight if your experience is elsewhere.
  • Comfort with data. You have pulled data out of Jira, a test-management tool or a database through an API or SQL, worked it in Python or Excel, and turned it into something a manager could read.
  • Writing that travels. Runbooks, skill definitions and findings clear enough to go to a client without you in the room.

What will decide the hire

  • Agency: a track record of picking up something nobody asked you to own and finishing it; a full day's progress without a stand-up; blockers surfaced early, in writing.
  • Curiosity and learning speed: when the ticket needs something you have not done, you learn it, try it, and come back with a working example rather than a question.
  • Opinions, held lightly: you form a view on how an agent or a pipeline should be built, say so, and change it when the evidence says otherwise.
  • Closure: working mechanism, a number that shows what changed, a written pattern, effort recorded.
  • Integrity with data: comfortable under rules that keep client data inside client systems and forbid attributing a problem to a supplier or a person.

What will help

Any two or three: a lived migration of agents between vendors, or hands-on time on a second enterprise platform (Gemini Enterprise, Vertex AI, Atlassian Rovo or similar); agent memory design (session, short-term, long-term, institutional) and the rules that stop it bloating; evaluating agent output (acceptance rates, LLM-as-judge, rubric scoring); retrieval-grounded agents over wiki content with source citation; MCP servers and tool integration; multi-agent or orchestrator patterns; contract testing (Pact) and service virtualisation (WireMock, Mountebank); performance scripting (JMeter); a dashboard tool (Power BI, Grafana); a test-management tool used through its API (PractiTest, Xray, Zephyr).

How we select

  1. CV review and fact-check: employers, dates and tooling claims are verified against public records and release dates. Dated, role-specific evidence beats a long tool list.
  2. Technical conversation, 45 minutes: one real agent run end to end, one real pipeline, one real data extract, and what went wrong in each.
  3. Practical exercise, one paid day: build a requirements-analysis agent on a synthetic requirements set we provide, with skills and knowledge structured for reuse and a short report on how often its findings were right; then wire its run into a CI pipeline with a notification. Working result, measure and write-up assessed equally.
  4. Engineering lead conversation: API and integration testing depth, CI/CD design decisions, hand-over.
  5. References and exclusivity confirmation.

To apply: send a CV with dated, role-specific tooling, and one paragraph describing an agent or workflow you built that nobody asked you to build, who used it, what it changed, and what you would do differently now.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company