Use case: Creating an agentic workflow from frontend to test automation
DEV Community

Use case: Creating an agentic workflow from frontend to test automation

Introduction

Every time our Frontend (FE) changed a UI component, a little nightmare started: tests broke, the Test Automation (TA) team was called and what seemed to be endless back-and-forth happened. I saw an opportunity to reduce this friction by introducing a headless agent into our workflow - one that picks up the broken tests and tries to fix them. Developing automations like this is a hard process though - designing a meaningful flow, dealing with cross-dependencies and accesses that might turn into dead-ends, and other unknown issues along the way. In this article, I'll walk you through the problem, the design decisions and trade-offs, and how it was implemented end-to-end.

The Problem

When working in frontend, we have some CI steps set up. One of them is to run test automation. More often than not, these tests fail because frontend changed the HTML structure or user flow altogether. When this happens, we have to go through these steps shown in the diagram.

Summarizing:

  1. FE: makes a breaking change.
  2. Tests break.
  3. TA: skip related tests.
  4. FE: CI pipeline now passes and merges the feature branch.
  5. TA: unskip the test, fix them, open a pull request and merge.

This workflow triggers for every feature branch opened by FE. The steps have to be done in sequence, and the flow involves at least four engineers (the ones writing the code and the ones reviewing it), context switching, and cross-team dependency. Also, there is a period in between that we lose test coverage: the tests pass, but only because they're skipped.

For our experiment, we decided to narrow our focus to something specific and easy to fix: when FE changes a UI component, needing a selector change; and only smoke tests would be used as the gate. I'll give you more context below.

What are selectors in UI testing?

In frontend applications, we implement components that users interact with. These components are accessible - either by their text, input placeholder, role or by non-visible screen-reader aria-* attributes, falling back to data-testid. These identifiers, selectors (or locators, as some libraries call them), are used by testing libraries, such as Testing Library, Playwright or Cypress to conduct automated tests, which simulate user flows.

Whenever a text, label, or whatever is used to identify the testing component is changed, the tests break. You may think this is a rare problem, but it can happen in some situations:

  • If the frontend is using a design system and it's moving to another one, this will likely break selectors. This will happen a lot during migrations.
  • If the text is changed and sometimes its position, this will likely break as well.
  • Any HTML change can break the selector.

What are smoke tests?

Smoke tests are a small set of end-to-end tests that cover only the most critical flows of an application. The idea is not to test everything, but to quickly answer one question: is the app still working? Because they are fewer and faster than the full suite, they are cheaper to run on every feature branch - and for this experiment, that also meant fewer (and less flaky) failures for the agent to deal with. If you want to go deeper, this article is a good starting point.

Constraints

This use case is not a one-size-fits-all solution, but it can give you some ideas. We had this specific setup:

  • Two separate codebases: some projects keep the frontend and test automation suite in the same codebase. This is not the case for our project - there is one repository for the frontend and another for test automation.
  • FE can't be built locally in TA CI: this means that the TA pipeline can't test on the same conditions as FE does. FE runs on a local build and triggers smoke tests there. On the other hand, TA can't - it can only run against some deployed QA environments. Bottom line: because of this, the same tests that fail in the FE pipeline can pass in TA (and vice-versa). This is why one of the agent's outcomes is "Not reproduced" (more on that below).
  • Different CI/CD tools per repository: FE runs its CI pipeline in Drone and its CD (deployments) in GitHub Actions (GA). TA only has a CI pipeline in Drone - it doesn't deploy anything, just some reports, which are managed by Drone as well.

Proposed Solution

The overall idea is that whenever there is a change in FE, the CI pipeline runs smoke tests, and if they fail, it would trigger a workflow agent in TA by sending a repository_dispatch event to the TA repo (through the GitHub REST API), which starts a GitHub Actions workflow there. This agent would then try to fix the non-passing tests. If it is successful, it would create a PR for a human to check. If this PR was merged, then the FE CI pipeline would pass.

Some notes to consider here:

  • When TA PR is merged, FE CI is not triggered automatically. A user has to re-trigger the job in Drone. This could be tightened in the future, by adding another task to do it when the branch is merged to main in TA.
  • Also when TA PR is merged to main, it will cause FE main and other branches to fail the tests (because they would not up-to-date yet with the feature branch). This is a known trade-off, which is acceptable, considering we are speeding up the process, at the same time controlling the flow (because a human will need to push the button to merge TA PR into main).

After coming to this high-level design (after a lot of back and forth), I started to think about the details.

Designing the automation flow

To implement this, some questions had to be answered:

  • Timing: FE can trigger the agent every time tests fail, but should it?
  • Platform: Where should the agentic workflow run - Drone or GitHub Actions?
  • Cross-repo/CI communication: How would the connection between FE and TA repositories happen?
  • Outcomes: What happens if the agent can't fix the test? What if it can? What if it can't reproduce (tests pass)?

Timing: when should the agent run?

Considering that every feature branch in FE triggers a CI pipeline and because of cost-efficiency and having more control, especially during the testing of this automation, I chose to trigger the agent under certain conditions. The feature branch changes would need to be deployed into a specific environment (we have more than 4 environments set up, so we can test multiple feature branches at the same time). Let's call this environment the QA Env. Having this environment is also needed because of the constraint that TA can only run smoke tests against deployed environments.

Platform: where should the agent run?

There are multiple steps of the automation design, and one of them is to decide where to trigger the agentic flow. As mentioned in the Constraints, we were using both Drone and GA. Our Drone was becoming crowded already - it was running processes on all branches, including main. And I realised that with GA we would have more control on when and how to trigger - their interface was more friendly for testing and manual triggering. Also, I could create a separated flow only for the agent, not mixing with other processes. For those reasons, I decided to go with GA for the agentic workflow, therefore the point of connection should be GA for triggering the agent.

Connecting the dots: how do the repositories communicate?

The initiator of this automation is FE when two conditions are met: smoke tests fail AND it is deployed to QA Env. Since FE CI runs in Drone and FE CD runs in GA, these two conditions are checked in different places. I had two options:

  1. Run smoke tests in GA after QA Env deployment, and then send the dispatch event to TA.
  2. In Drone, after smoke tests step, check if QA Env is deployed and then send the dispatch event to TA.

The first option would be more costly - we would need to bring the smoke test step into the GA, and it would duplicate the step in both CI and CD pipelines. The second option, on the other hand, would be a simple check. That's because we use GitHub labels to track which environment is being deployed. We already had a flow set up where adding a label to a PR triggers a deployment to that environment. Given that, we could check the GitHub label right after the Smoke tests step. If both smoke tests failed and the deployment label was set as QA Env, then Drone would send the dispatch event to TA.

Given that CI and CD happen in different environments, one trade-off here is that it is possible that Drone will receive and identify the label, but the GA deployment process can fail, or Drone check can happen before the branch is actually deployed. I accepted these trade-offs, because the worst that could happen is that the tests run against an environment without FE change, triggering a false negative. For the race condition, as soon as the label is added, usually the CD happens in 5 minutes. The entire CI pipeline usually takes 15 minutes, making it a comfortable pass, since the check would be as the last step of the CI pipeline.

Outcomes: what should the agent report?

The FE created a new feature that broke smoke tests, it was deployed to QA Env, and the dispatch event was sent. Now what? What should the agent do? The entire goal is for the broken smoke tests to be fixed, by updating the selectors. For that, I took some decisions:

We had more than 40 written smoke tests in total. For the purposes of this automation, we don't need the entire suite to run. Therefore, there should be a payload carrying the information of which smoke tests have failed in the FE side. When attempting to fix the test, I see some outcomes, overall:

  • The agent is able to fix the test.
  • The agent is not able to fix the test.
  • The tests pass.

So, with three main outcomes, plus two edge cases we would have:

  • If tests are fixed, then a PR is created in TA side, and a comment is sent to FE PR with the message of "Already repaired", along with the link to the TA PR.
  • If tests were already fixed by a prior run, then a comment is sent to FE PR with the message of "Already repaired", without a link to a TA PR.
  • If tests are not fixed, then no PR is created, but a message is sent back to FE PR with the message of "Possible regression".
  • If tests pass, then a message is sent back to FE PR with the message of "Not reproduced".
  • If there is an environment failure (e.g., QA Env is unreachable), the agent aborts and no message is sent back to FE PR - the error is only logged in the GA run.

Although it is possible to see all the logs in GA workflows page, I wanted to make provide as much up front transparency as possible, by ending the flow in FE PR, with a message on what was the outcome of the entire process. Then, with most of the definitions set up, we can finally look at how the agent does this, and the harnesses around it, in the next sections.

Setting up the prerequisites

As you will see, there are multiple steps to implement the flow end-to-end.

Creating the agent harness

A harness is everything around the model that shapes how the agent works - instruction files, skills, commands, scripts and guardrails - so it behaves predictably without someone guiding it step by step. The initial state of the codebase was not ready for AI at all. To avoid increasing the blast radius of the project, I decided to just create enough harnesses to implement this POC. Two assets were created:

  • AGENTS.md : with CLAUDE.md pointing to it, which enables uses for both Claude and OpenAI models to read. Normal instructions file that is injected to every agent session (not going too deep into this as there's plenty of documentation on the web).
  • Selector conventions skill : defines the guidelines on how to choose the selector. There is a hierarchy the agent has to follow. For example, the first attempt should be with getByRole, then by getByLabel and so on.

With those, I tested how it performed using my own agent (local Claude Code CLI). Other commands were created too, but they belong to the automation per se, so they will be specified in the following sections.

Making tests environment-agnostic

Another problem was that the tests were not able to run against QA Env. Data was coupled into a specific environment, as well as some selectors. This isn't related to AI nor to automation, but it would impact the ability to develop it. Therefore, I had an additional prep work: creating functions that would setup/create and teardown/delete data for every test run and making selectors environment agnostic.

Implementing the flow

A visual diagram of the entire flow is shown below:

Frontend: dispatching the event

As the flow starts in FE, naturally I had to start there, but the work was slim: I've added a new step in drone.yml, after the Smoke test step:

- name: Smoke test
  commands: # Smoke tests instructions
- name: Dispatch selector repair environment
  GITHUB_APP_TOKEN:
    from_secret: GITHUB_APP_TOKEN
  REPAIR_ENVIRONMENTS: QA_Env
  commands:
    - sh ci/drone/dispatch-selector-repair.sh
  depends_on:
    - Smoke test
  when:
    status:
      - failure

(Note: this is an example piece of code and should not be used as it will not work).

As variables, a GitHub App token was needed to enable cross-repository communication and the REPAIR_ENVIRONMENTS is a mirror list of which environments are enabled. For now, only QA Env is enabled to trigger this automation.

I could add all the instructions in the YAML file, but I had extracted them into ci/drone/dispatch-selector-repair.sh to separate concerns and improve readability. In this script, a few things were done:

  • Check if deployed environment is valid
  • Check if GitHub token is valid
  • Collect the failed tests. Exit if no JSON report of failed tests is found.
  • Fetch commit SHA, PR id number
  • Prepare payload and dispatch frontend-smoke-failed event

Then, TA would receive this event and payload in its end.

Test Automation: running the agent

To receive the event, a GA workflow was created. selector-repair.yml is triggered when the event is received:

repository_dispatch:
  types: [ frontend-smoke-failed ]

It can also be triggered manually, although this was mostly for testing purposes.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.