CUEDEV
← All field notes
Controls Sep 2026 12 min read

When Are SOX 404(b) Controls Tested During the Year?

For most public companies, SOX testing is not random. It follows a yearly rhythm that matches the finance team's actual pain: plan in spring, document in early summer, then test in waves around every close. This is that calendar, month by month, plus exactly where agentic AI automation plugs in to watch the controls, collect the proof, and cut the grind.

6
Calendar stops, April to February
5
Value-table comparisons
50-70%
Evidence collection time cut
5
FAQs answered
01

Year at a glance: the SOX testing calendar

Take a 12/31 year-end company. The rough pattern looks like this:

  1. April - plan the SOX year
  2. May - update process and control documentation
  3. June - walkthroughs and test of design
  4. Aug-Sep - test controls for Q1 and Q2
  5. Nov-early Dec - test controls for Q3
  6. February (next year) - test Q4 activity and finish

Every gap in that list is a month where finance is buried in close, 10-Q, or year-end work. SOX testing has to live around that.

An AI agent can sit under this whole year, tracking control runs, logging evidence, and flagging odd items in real time, so when humans arrive to test, they are reviewing, not hunting.
02

April: planning the year without burning out

April is the flight-plan month. You answer which controls are in scope, how many samples per control, who owns what, and what the testing calendar looks like for each control family. Right now that lives in spreadsheets, random slides, and long email threads.

Real story

One controller we worked with had three different SOX calendars for the same year. Audit, internal control, and FP&A all had their own. They were all right and all different.

An AI agent takes your scoping rules and creates a single, living SOX calendar. It can pull last year's controls, owners, and issues, map them to this year's org and systems, auto-create a draft test plan by quarter, and keep that plan in sync as owners or systems change. That shifts April from version drama to review-a-draft-and-tweak, and can cut planning time by 40 to 60 percent.

03

May: documentation clean-up and change tracking

May is where you fix the story: updating narratives for process owner changes, capturing tweaks to controls, aligning flowcharts with how people actually work now. The real problem is that someone changes a step in October, nobody updates the narrative, and by May nobody remembers why it changed.

An AI agent watches the systems all year and proposes document updates in May. It compares last year's control steps to this year's activity, spots new approvers or extra review layers, drafts redlines for narratives and RCMs, and flags stealth controls that were added or dropped along the way.

May stops being interview-twelve-busy-people-and-guess. It becomes review-what-the-agent-saw-in-the-logs-and-click-accept-or-reject, cutting document refresh time by 50 to 70 percent and making the docs match real life instead of wishful thinking.

04

June: walkthroughs and test of one

June is walkthrough month. A walkthrough means picking one transaction, following it end to end, and proving the control design actually covers the risk, exactly the procedure PCAOB AS 2201 ↗ calls out as ordinarily sufficient to evaluate design effectiveness. You are not testing volume here. You are testing design, the same discipline we walk through in 4 steps to design internal controls.

The pain point is that finding a good example, getting all the screens, and stitching the story together is slow. An AI agent can auto-build full end-to-end walkthrough packs: pick a clean, typical transaction, pull screenshots, logs, approvals, and timestamps, lay them out in order with short notes, and highlight where key controls fire.

You still do the judgment. But the hunting is gone.

Walkthrough prep time can drop from days to hours, and design gaps show up earlier, while there is still time to fix them.

05

Aug-Sep: testing Q1 and Q2 activity

After Q2 close, the team can breathe, and that is when testing of operating effectiveness for the first half kicks in: pulling samples from Q1 and Q2 to prove controls actually ran, not just that they exist on paper. Manual life right now means pulling a population from the ERP or sub-systems, cleaning the data, drawing samples, and chasing evidence across email, shared drives, and ticket tools.

An AI agent can build the population in real time and pre-attach evidence all year: tracking each time a control fires in production, tagging the event with control ID, owner, time, and system, and storing the key evidence instantly. The same completeness-and-accuracy discipline behind our IPE testing template applies directly here, just running continuously instead of once a quarter.

Sampling turns into click-to-sample-and-review rather than beg-IT-for-a-dump-and-clean-it-for-a-week. Time saved for Q1/Q2 testing is often 50 percent or more, and error rates in populations drop sharply because the agent sits closer to the source.

06

Nov-Dec: testing Q3 activity

By November, Q3 close is over. Now you test controls for July, August, and September, and check that fixes from earlier issues are really working, right as people are already tired and starting to think about budgets and year-end.

The agent turns Q3 testing into a delta check rather than a full dig: comparing Q3 control runs to Q1/Q2 patterns, spotting drifts like late approvals or odd timing, pre-flagging items that are likely fails, and showing, for each control, what percentage of runs met all the rules.

Instead of looking at every sample with the same energy, humans zero in on the 10 to 20 percent that actually look risky. That shrinks test effort and improves coverage at the same time.

07

February: testing Q4 and closing the loop

January is pure chaos: year-end, audit, disclosures. So Q4 SOX testing lands in February, covering October through December, confirming yearly controls ran, and finalizing the Section 404(b) ↗ conclusion for the 10-K. Two classic headaches show up here: yearly controls where proof gets lost, and late Q4 surprises that surface just before filing.

The agent babysits the year-end controls as they run and checks completeness in real time: watching for yearly controls that have not fired by a set date, pinging owners with clear reminders, capturing proof the moment it runs, and checking that every key Q4 control has at least one clean sample. By February you are reviewing a dashboard of facts instead of guessing, which cuts the last-minute scramble and lowers the odds of a material-weakness surprise. The same period-long conclusion shows up in a SOC 2 Type II report, covered in SOC 1 vs SOC 2 vs Type I vs Type II.

08

How AI agentic automation fits into this calendar

The classic SOX year is long quiet stretches where controls quietly run, broken up by short intense bursts where humans rush to prove they ran. Traditional SOX tooling helps with checklists and storage, but it is still human-pull: you go in, click, upload, and reconcile. Agentic AI flips it to system-push.

  • Instead of collecting evidence once a quarter, agents collect it every time a control runs
  • Instead of updating docs once a year, agents watch changes and propose redlines
  • Instead of designing a test plan once a quarter, agents roll the plan forward and adjust in real time

The idea behind CueDev's work here is simple: plug agents into the tools you already use (ERP, ticketing, email, close tools), teach them your control set and risk logic, and let them track, summarize, and flag while humans stay on the high-judgment work. That gives you a SOX program that is faster, cheaper, and far more ready when an auditor starts asking to see it.

09

Quick value table

Here is where the gains usually show up:

AreaTypical manual realityWith agentic automation
SOX planning effort3-4 weeks of scattered work each April1-2 weeks, 40-60% less manual setup
Evidence collection time60-70% of tester hours stuck finding filesCut 50-70%, evidence pre-linked to events
Documentation refreshFull-month scramble every May50-70% faster via auto-drafted redlines
Sample selection accuracyManual Excel filters and one-off queriesClean, system-level populations and smart samples
Last-minute year-end fixesFrequent fire drills in Jan-Feb30-50% fewer late surprises from missing runs

The exact numbers shift by company, but the pattern is stable: most gains come from killing the hunt, not the thinking.

10

FAQ

Because risks do not wait for December. Testing in waves (Q1/Q2, then Q3, then Q4) lets you spot issues early and give management time to fix them before they turn into a year-end problem. Agentic AI helps by watching all year, not just in test months, so those issues show up even faster.

The June walkthrough is where you prove the design of a control makes sense. If the design is weak, testing more samples later only proves the same thing many times. Using an AI agent to build clear walkthrough packs catches design gaps early and avoids rework across the rest of the year.

Auditors care about source reliability, clear linkage from control to event to proof, and the ability to re-perform. Agentic automation collects from the same systems auditors already rely on and stores a clean trail of who did what, when, and where the data came from, which often reduces back-and-forth as long as the setup is transparent and well documented.

The calendar here fits a classic 12/31 filer under 404(b), but the pattern is just as useful for smaller issuers, pre-IPO companies preparing for control reviews, and private firms under bank or board pressure for stronger controls. The same agentic setup scales down: fewer controls, same logic, less noise.

Not from AI writing your 10-K. It comes from killing duplicate data pulls, standardizing how evidence is named and stored, cutting down on resend-that-PDF emails, and letting humans spend most of their time on exceptions instead of routine passes. That is where firms typically see 10-30% lower outside audit time and 30-60% less internal manual effort.

11

Next step

If your SOX calendar today feels like a yearly stress test, you do not need a bigger spreadsheet. You need a quiet layer of agentic automation living underneath the year, so testing months become review months, not rescue missions.

Book an audit →