An agent that misreads a VIN, approves coverage on an excluded part, or grants a discount nobody authorized isn’t a UX bug. In a dealership or a plant, it’s a chargeback, a warranty dispute, or a margin leak that shows up on someone’s P&L weeks later. That’s the risk profile you’re actually testing against - not “does the bot sound helpful.”
The architecture has to split the work
Salesforce’s Atlas Reasoning Engine plans conversations the way a person would: read intent, pull context, sequence actions, reflect, respond. Left alone, that probabilistic loop is exactly the wrong tool for warranty adjudication and pricing authority ; a slightly different phrasing shouldn’t produce a different financial outcome, and auditors don’t accept “the model felt like it that day” as a control.
Agent Script exists to draw that line. The generative model still owns tone and comprehension; Agent Script owns state, conditional logic, and which actions are even visible to the reasoning engine at a given moment. Testing has to check both halves that the agent still understands the customer, and that the deterministic rails around it haven’t been quietly widened.
Three exceptions worth testing before go-live
Missing or superseded parts. Parts catalogs break in predictable ways - a part is discontinued, superseded twice, or split by a chassis or serial-number cutoff. A well-built agent doesn’t guess. Tested correctly, it traces the supersession chain to the active SKU, checks fitment, and when a fitment split exists, stops and asks for the serial number instead of shipping the wrong bracket. The failure mode worth catching in testing isn’t a wrong answer; it’s a confident one built on a part number that no longer exists.
Disputed warranty coverage. This is where agents are easiest to over-trust. A customer reports a failure at 58,000 miles inside a 60,000-mile warranty; telematics shows the fault code logged at 61,200. The agent’s job is to surface that discrepancy, explain the coverage terms, and route the claim and never to say “this is covered.” Conversation tests here are really guardrail tests: assert that no agent response contains a binding commitment, and that ambiguous cases (wear-and-tear versus defect, an active safety recall, a non-OEM modification) get routed to a human rather than settled in-chat.
Approval exceptions. A loyal customer with ten machines asks you to waive a $3,000 repair bill on the spot. The correct behavior isn’t a diplomatic “no” generated by the model; it’s that the discount and goodwill actions are architecturally invisible to the agent until an approval-tier variable is set, at which point it opens a request to a supervisor instead of negotiating. Test suites should specifically try to talk the agent past its own authority limit and confirm the ceiling holds regardless of how the request is phrased (“I’ll cancel the contract,” “I’ve bought ten machines from you”).
Why this has to run continuously, not once
Parts inventories, warranty policies, and the underlying models all change on their own schedules, so a pre-launch test pass is a snapshot, not a guarantee. The discipline that holds up is treating conversation tests as CI/CD regression gates - a pull request that touches Agent Script or an invocable action triggers a scratch-org run of the full suite before anything reaches staging.
Salesforce’s own testing documentation gives a sense of the scale this is built for: the Testing Center supports up to 1,000 test cases per run and up to 10 test jobs in a 10-hour window, which is the difference between “we checked it once” and “we can regression-test every change.” When a live conversation ends in a negative rating or a human takeover, that transcript becomes tomorrow’s new synthetic test case - the suite grows from the exceptions the business actually hits, not just the ones a team anticipated in advance.
The synthesis
None of this is really about the model getting smarter. It’s about deciding, in advance and in writing, what an agent is allowed to promise, and building the tests that prove it stays inside that line on the parts counter, on a disputed claim, and on a discount request that arrives with real pressure behind it.
If you’re taking an agent live in a dealer or plant environment: which exception would embarrass you first if it slipped through untested - a parts fitment error, a warranty overreach, or an unauthorized discount?
Ferfier builds these conversation-level test suites for Agentforce deployments, treating every parts, warranty, and approval exception as a regression gate rather than a one-time launch check. If you want to see which exception would embarrass your own agent first, talk to us.