Loading presentation...
How Cucumber scenarios make code legible to AI agents
Out-of-stock products, API timeouts, and catalog sync delays cause template resolution to fail. Without filtering, raw placeholders leak to the user.
// Old bestAvailableFallback — no filtering private List<OrderConfirmation> bestAvailableConfirmation( String channel, String productType, String locale, int maxItems) { List<OrderConfirmation> confirmations = templateProvider.getConfirmations( channel, productType, locale, maxItems); if (confirmations.size() >= maxItems) return confirmations; // {product_name} leaks here! }
The agent has to trace through template loading, product resolution, and exception handlers to understand what can go wrong
Scenario: Out-of-stock product shows clean confirmation Given the product type is "ELECTRONICS" And the product ID is "SKU-12345" When the catalog service returns no information And order confirmations are requested Then all confirmations should not contain any placeholder pattern
The agent grasps the invariant instantly: when the catalog fails, no placeholders leak to the user
Scenario: Catalog succeeds and templates are properly resolved Given the product type is "ELECTRONICS" And the product ID is "SKU-67890" When the catalog service returns "MacBook Pro" as the product name And order confirmations are requested Then some confirmations should contain "MacBook Pro" And no confirmations should contain "{product_name}" And all confirmations should have resolved text
Happy path and safety check in one scenario: templates resolve AND no placeholders leak
One Gherkin block, 10 test executions. The agent sees every combination at a glance.
Scenario: All locales have product confirmation supply Given the product type is "<productType>" And the locale is "<locale>" When the catalog service returns no information And order confirmations are requested Then some confirmations should contain a subject line And some confirmations should contain shipping details And all confirmations should not contain any placeholder pattern Examples: | productType | locale | | ELECTRONICS | en-us | | CLOTHING | en-us | | ELECTRONICS | de-de | | CLOTHING | de-de | | ELECTRONICS | ja-jp | | CLOTHING | ja-jp | | ...4 more rows |
Scenario Outline: All types stay clean Given the product type is "<productType>" When catalog returns no information Then no placeholder pattern Examples: | productType | | ELECTRONICS | | CLOTHING | | BOOKS | | FOOD |
Scenario: NONE confirmations Given the product type is "NONE" When order confirmations requested Then no placeholder pattern
Scenario Outline runs the same test across every product type. One Gherkin block, four test executions.
PR review says "use three confidence layers." Write the scenarios first, agree on the wording, then implement.
Scenario: Express shipping when confidence is HIGH Given a shipping option with text "delivers to {address}" When the shipping provider recommends "123 Main St" with HIGH confidence Then a resolved option "delivers to 123 Main St" should be present Scenario: Generic fallback when confidence is LOW Given a shipping option with text "delivers to {address}" And a shipping option with text "standard shipping available" When the shipping provider recommends "123 Main St" with LOW confidence Then a resolved option "standard shipping available" should be present And no resolved option should contain "123 Main St" Scenario: No option when no recommendation exists When the shipping provider has no recommendation Then no resolved options should be returned
The original placeholder fix seeded a living spec across 8 feature files
Plain text Cucumber scenarios are the most agent-readable format. No AST parsing, no class hierarchy traversal.
Every use case is an executable spec. Agents know exactly what the system does and what edge cases are handled.
Gherkin is the common language between humans, agents, and CI. Everyone reads the same source of truth.
The service ships both CLAUDE.md and AGENTS.md with the same BDD instructions. Any agent that opens the repo learns the rules.
A new agent reads CLAUDE.md to learn the workflow, then reads .feature files to learn every behavior. Day one contributor.
Scenarios are cheap. Implementation plans are not. Agree on behavior first.
The same code that's hard for humans to parse at a glance is even harder for agents to reason about
// order-service // 2026█
Everything you need to add BDD to a backend service.
Start with one feature file for your most error-prone path. The framework pays for itself on the second bug fix.