The Claim, and Its Boundary
An autonomous SEO system does not fail only because a language model writes a weak answer. It fails when an observation is mistaken for a current fact, when a query is mistaken for page intent, when two agents receive different strategic truths, when a correct change is applied twice, when a process dies after publishing but before recording the receipt, or when a report converts missing evidence into a persuasive zero. Model capability matters. The state surrounding the model determines whether that capability can produce a reliable result.
Our central systems hypothesis is that agent success becomes more probable when a system controls six surfaces explicitly—epistemic state, attention, strategic authority, temporal validity, operational side effects, and human governance. This is consistent with research showing that reasoning-and-action loops, retrieval, decomposition, tool use, and feedback can extend language-model behavior beyond one-shot generation.[4][5][7][8][9] It is also consistent with evidence that simply increasing context length does not guarantee that a model will use the right information, especially when relevant material sits in the middle of a long input.[6]
These studies motivate individual components; none evaluates their joint effect in this SEO system.
The boundary is equally important. We have not yet run a controlled before/after benchmark showing that this architecture increases autonomous task success by a measured percentage. We can show implemented mechanisms, contract tests, dated production observations, and dated failures those mechanisms prevent. We can show search growth that motivated the program. We cannot honestly call that growth a causal estimate of the new architecture. The contribution of this paper is a falsifiable systems hypothesis and an implementation case study, not a randomized trial.
Claim discipline
The architecture makes identifiable failure channels explicit and is designed to reduce them. The magnitude of any resulting uplift remains an empirical question. Wherever this paper says “advantage,” it means a mechanism that makes a necessary success condition explicit, testable, recoverable, or reviewable—not an unmeasured percentage improvement.
The Founding Observation: What 12× Means
This paper is anchored by an observational founding result. On an anonymized current-events property, the first 28 days of Search Console data averaged 122.8 impressions per day. The latest 28 complete days at capture averaged 2,101.9 per day: a ratio of 17.1. The public phrase “more than 12×” was deliberately conservative.[1] A 12× outcome means the later value is 1,200% of the baseline, or a 1,100% relative increase; it does not mean a 12% improvement. The observed 17.1× outcome means roughly 1,710% of baseline, or a 1,610% relative increase.
The calculation is intentionally window-to-window rather than peak-to-trough. Fixed 28-day windows balance day-of-week composition and make the denominator inspectable. Search Console is still an instrument with aggregation, privacy, and data-limit behaviors, not an omniscient record.[2] The property is demand-volatile, the available founding set is two sites, impressions are visibility rather than revenue, and the repository record is a curated series plus a captured screenshot rather than an independently archived raw export. The exact daily rows and fixed-window boundaries are not retained in the public repository, and the chart's ISO-week boundaries therefore cannot independently reproduce the reported daily-window ratio. A second local-service property reached an observed impressions plateau, illustrating why local-grid coverage can be more decision-relevant for that case than a rising impressions curve.
Most importantly, this observation predates the complete closed-loop architecture described below. It is founding evidence: it showed that an operating pipeline could coincide with a sustained visibility shift over the observed period and exposed the need for a system that could determine what to do next. It is not causal evidence that the strategy compiler created a 17.1× outcome. A causal claim would require a prespecified counterfactual design, stable covariates, and an explicit treatment date—for example, a Bayesian structural time-series analysis under its stated assumptions.[3]
A Closed-Loop Model of SEO Work
The first pipeline was largely feed-forward: collect signals, generate work, publish, and repeat. The current design is a feedback controller at the systems-architecture level. Its state is not hidden in a prompt; it is represented as evidence, an active strategy version, compiled job context, durable execution state, and measured outcomes. The controller changes policy through a governed reconciliation workflow that culminates, when policy permits, in transactional activation, rather than allowing a model to rewrite its own instructions. We borrow the observation–decision–action–feedback vocabulary of control systems, but do not claim classical stability, controllability, or observability guarantees for the SEO environment.[27]
Current producers populate subsets of this target. Research captures store provider, source kind, request fingerprint, dimensions, observation time, effective period, expiry, raw reference or hash, compact payload, and classification; family-specific caches retain other fields. The tuple is a design target, not a claim that every legacy row has every coordinate.
This separates three concepts that are often collapsed. Evidence is an observation with scope and time semantics. Strategy is the current policy for interpreting those observations and assigning ownership. Execution is a bounded attempt to change the world under that policy. An outcome becomes a durable reconciliation input—and, where a receipt producer mints it, new evidence—for a later transition; it does not retroactively redefine the policy that authorized it.
- 01
Observe
Search, analytics, rank, local, speed, crawl, schema, inventory, and research signals.
- 02
Persist
Project-scoped caches retain provider data; research captures additionally mint compact receipts with observation and provenance fields.
- 03
Qualify
Where adopted, freshness, comparability, scope, and provenance determine decision eligibility; legacy consumers remain outside this guarantee.
- 04
Compile Strategy
A versioned active policy binds clusters, page intent, priorities, risks, and measurement.
- 05
Project Context
Compiled-context jobs receive a bounded strategy, evidence, and site-rule projection; other lanes consume projected artifacts or typed prompts.
- 06
Execute and Verify
Agent lanes use their applicable durable dispatch or job, bounded tools, and validation; publication work additionally receives Git and deployment receipts.
- 07
Measure and Reconcile
Outcomes, errors, reports, and human review feed reconciliation under their own provenance and lifecycle rules.
The distinction matters because SEO is partially observable and delayed. Search demand, crawling, indexing, ranking, local visibility, engagement, site deployment, and competitor behavior update on different clocks. A closed loop cannot eliminate that uncertainty. It can preserve which clock a claim came from and refuse to compare windows that do not mean the same thing.
The Evidence Plane
The evidence plane begins before the model. Google Search Console contributes query-level impressions, clicks, click-through rate, and position. GA4 contributes sessions, users, and acquisition behavior; Google itself cautions that Search Console and Analytics measure different events and should be interpreted together rather than forced into numerical equality.[25] DataForSEO contributes keyword, result-page, competitor, authority, and “People also ask” observations. Geographic grid scans measure local results from explicit coordinates. PageSpeed, crawl, sitemap, schema, index checks, live-site probes, and the site inventory describe technical state. GitHub, hosting, and live-check receipts describe what was actually shipped.[19]
The target evidence contract asks for more than a value: tenant, project, URL or query, country, language, device, location, effective period, observation time, method, completeness, and source. Research captures mint compact receipts separately from large provider caches; other provider families retain their own subsets. Paid operations receive request fingerprints, and a source-health ledger records operational state. This follows the same intellectual direction as formal provenance models: an assertion should retain the entity, activity, and source relationships that made it possible.[12][20]
Measured
Direct provider or live-system observation with a named window and method.
Cached
Previously measured evidence reused inside an explicit freshness policy.
Inferred
A derived conclusion whose inputs and method remain inspectable.
Echo
A repeated presentation of an older observation, not a new measurement.
Unavailable
No trustworthy observation exists; absence stays absence rather than zero.
Stale
In migrated consumers, remains visible for diagnosis while carrying zero decision weight.
Where the shared eligibility contract has been adopted, these states prevent three common category errors. First, a cached rank row is not a new rank measurement. Second, missing analytics is not zero traffic. Third, a report window is not comparable to another merely because both contain a number. The selector can keep stale evidence visible for diagnosis while assigning it zero decision weight. Paid weekly rank observations outrank daily cache echoes on the same day; location and device identity travel with the observation; report deltas name their windows.[20]
A current limitation
Freshness eligibility is centralized, but it is not yet universal. Strategy review and FAQ selection use strict local rules; some legacy article, recommendation, radar, and autoblog paths still consume projected or cached artifacts without the same complete provenance envelope. The system is moving toward one evidence contract, but this paper does not claim that every consumer has arrived there.
Strategy as a Versioned Compiler
A strategy is not a dashboard summary and not a prose document that each agent interprets independently. It is a versioned authority with a content hash. The research pipeline assembles a draft through named stages, validates structured synthesis against evidence IDs, and compiles a preview. Activation is transactional: one version becomes active; its backlog remains versioned strategy state, while page intents, page-keyword ownership, tracked terms, and eligible recommendation-channel items are projected into operational tables.[19][21]
| Stage | Responsibility |
|---|---|
| S0 — Intake | Create or reuse the draft shell; combine services, locations, seed terms, and observed queries. |
| S1 — Demand | Build and deterministically pre-cluster the keyword universe and its demand signals. |
| S2 — SERP & competitors | Measure search-result composition, ranked competitors, authority gaps, and local-pack conditions. |
| S3 — Entities & content | Inspect topical entities and the content of pages Google currently rewards. |
| S4 — Inventory | Reconcile the live site, canonical paths, internal structure, schema, crawl, and index state. |
| S5 — Baselines | Assemble per-cluster 28-day Search Console and latest rank baselines before synthesis. |
| S6a–c — Synthesis | Produce evidence-cited clusters, an operating charter whose factual support remains on its children, and a typed Content/Page/Technical backlog. |
| S7 — Compile | Render the exact strategy preview and agent-facing form from the validated draft. |
| S8 — Reconcile | Assemble nine deterministic strategy, evidence, and state blocks; stamp the active content hash and input digest; propose a bounded amendment under the current non-uniform selection and eligibility policies. |
Authority has an explicit precedence. Human-authored page intent outranks strategy-derived intent; strategy-derived intent outranks automatic query observations. Search Console may refresh the queries associated with a page, but it may not convert a popular query into the page's primary intent or overwrite a human decision. Workstream is stored separately from execution channel: “Page” describes who owns a search-intent problem, while the same local executor can still implement both Page and Technical changes.[23]
Synthesis also separates model judgment from factual authority. The model can organize clusters, articulate a charter, and propose work, but synthesis does not mint provider evidence. Cluster and backlog outputs cite existing evidence IDs where receipts support the claim; explicitly synthesized or technical-health-only claims may use a closed synthesis code with no evidence ID, and the charter delegates citations to its children. Stored artifacts and hashes make the run auditable. The compiled projection is deterministic over stored strategy state apart from its compile timestamp; stochastic model synthesis is not claimed to be byte-for-byte reproducible. This is closer to a compiler pipeline than to a chat session: parse observations, type them, construct an intermediate representation, validate it, and emit projections.
Versioned Strategy + Content Hash
Evidence references and codes, page ownership, priorities, constraints, risks, and measurement policy.
Content
Topic selection, briefs, drafting, differentiation, publication cadence, and factual grounding.
Page
Search intent, keyword ownership, cannibalization, copy, metadata, visible FAQs, and links.
Technical
Crawl and index state, performance, technical schema, deployment, and off-site work.
Results
Comparable performance windows, local visibility, report QA, and client-facing evidence.
Reconciliation
Newly assembled strategy, evidence, and state blocks, outcomes, errors, and accepted findings propose the next strategy state.
Context Projection, Not Prompt Accumulation
Retrieval-augmented generation established that external knowledge can be brought into a model's working context rather than embedded permanently in model parameters.[5]Our operational problem is narrower: from a large client state, what is the minimum complete dossier for this job? More context is not automatically safer. Long-context research shows that position affects use; our own earlier incident showed that a correct file could exist in the workspace yet remain unread because the read-order index omitted it.[6][24]
Although the context-pack row contract admits blog-draft and apply-edit tasks, the current headless runtime materializes the full Context Pack V2 only for a blog-draft dispatch. When enabled, the compiler stores a fixed seven-file bundle capped at 48 KB: an index, task dossier, active strategy, guardrails, site framework, recent evidence, and live inventory. Protected-path and acceptance-criteria material is co-located in the task and guardrail sections. The terminal loads the immutable row by exact project, task, and dispatch identity and materializes it into the dispatch workspace with exclusive-create and anti-symlink checks. Apply-edit deliberately materializes only its framework document while its task block travels in the edge-composed prompt. Non-file worker jobs use a separate private scratch directory.[21]
The context invariant
Two agents working for the same project may receive different evidence slices, tools, and acceptance criteria. They must not receive competing active strategies. Specialization changes the projection; activation changes the authority.
This projection is designed to improve attentional conditions in three ways. It removes irrelevant provider payloads, places the job specification beside the strategy rules it depends on, and gives validators stable identifiers for the claims an output must satisfy. It also constrains cost: context is read on every turn, so an unbounded dossier consumes both attention and tokens. The pack is a budget allocation decision, not merely a prompt-format decision.
Direct consumption is intentionally not identical in every lane. Local blog-draft workspace agents, when their context path is enabled, and report strategy sections consume the active compiled strategy directly. Some in-edge content paths consume transactionally projected recommendations, opportunities, and gates instead of reading the full compiled pack. Context Pack V2 is also configuration-gated and defaults off in source unless the environment explicitly enables it. These are material boundaries: the compiler is the architectural authority, but not every model invocation presently receives the same physical artifact.
Typed Agents and Shared Execution
We use “multi-agent” to mean role-specialized execution over shared state, not a room of unconstrained personas debating one another. Research agents collect and synthesize. Content agents draft against a brief. Page-owned tasks resolve intent, differentiation, cannibalization, copy, metadata, visible FAQ, and links. Technical-owned tasks handle crawl, index, performance, schema mechanics, deployment, and off-site implementation. Reporting workflows freeze measured windows and client-visible evidence. A reconciliation agent proposes a bounded strategy amendment. Multi-agent frameworks demonstrate the composability of such roles; benchmarks also show that agent performance is environment- and task-dependent, which is why the orchestration contract matters as much as the cast of agents.[10][11]
The Page and Technical surfaces deliberately reuse one execution substrate. A single classifier assigns each item to an owner-facing workstream, but recommendations and live fixes converge on the same detail, validation, local edit, Git publication, deployment, and verification machinery. This avoids the failure mode in which duplicate dashboards create duplicate queues and contradictory terminal state.
- Content: local headless agents consume the active strategy directly; in-edge composition receives projected recommendations, opportunities, and gates, with factual-grounding, uniqueness, internal-link, image, QA, and publication rules.
- Page: cluster ownership, page intent, target queries, cannibalization, existing copy, visible FAQ, and internal links, then the shared fix executor.
- Technical: PageSpeed, crawl, sitemap, schema, index, probe, hosting, and repository evidence, then the same executor with technical validators.
- Results: frozen search and analytics windows, local-grid identity, report QA, attachment, delivery, and client-visibility gates.
- Strategy reconciliation: active content hash plus nine deterministic evidence blocks, strict finding vocabulary, authorship protection, and one successor if evidence changes mid-run.
The design favors workflows for predictable paths and agent discretion only where semantic judgment is necessary, matching the practical distinction between predefined workflows and model-directed agents.[15] Models choose among bounded semantic alternatives; databases choose ownership, authorization, idempotency, and terminal state.
Local-First, Durable Execution
Eligible SEO agent lanes prefer a bound desktop already paid for through a subscription. The server does not assume that a machine is eligible because it recently sent a heartbeat. The heartbeat advertises capability only with a proved subscription login, at least one caller-visible binding, and healthy callback persistence. Each claim then rechecks the local run permission, caller/tenant/project visibility, machine ownership, requested-project binding, and job-kind capability before its queued-to-running compare-and-swap and callback token rotation.[22]
For model-only worker jobs—topic refill, FAQ answer drafting, and strategy reconciliation—the runner invokes one-shot model execution in a private 0700 scratch directory, without a repository worktree or API key. Blog and fix lanes instead run against isolated or bound Git workspaces so they can return a bounded change set. No local SEO runner pushes; the edge publisher owns Git publication state and coordinates the applicable deploy and live-check follow-through. This split keeps local computation close to owned resources while preserving a server-authoritative publication boundary—a practical adaptation of local-first principles to an operational system rather than a collaborative document editor.[14]
- 1
Queued
Durable row; no paid spend merely for waiting.
- 2
Eligible
Tenant, user, project binding, permission, capability, and policy agree.
- 3
Claimed
Compare-and-swap winner rotates a single-purpose callback token.
- 4
Executing
A local subscription runner is preferred; an allowed hosted lane can fall back.
- 5
Finalizing
The lane validates output and records durable handoff state around its remote boundary according to its contract.
- 6
Completed
The applicable terminal receipt is recorded without treating skipped planes as completed.
Recovery branch
Stale claims, ambiguous responses, and process restarts reconcile from durable state. Retries carry stable or versioned attempt identity so receivers distinguish replay from authorized re-execution; ambiguous provider egress is conservatively accounted.
Publication branch
File-edit agents return bounded change sets; model-only workers return validated structured results. The Edge writer alone owns Git and pull requests, then coordinates the applicable deployment and live-proof receipt.
The product contains two separate durability systems. Builder Automations reserve local finalization capacity before the database claim, persist the result before optional auto-push, and replay completed records with a matching fresh tenant/user caller. SEO blog, fix, and worker jobs use the headless callback outbox plus database claim/finalize state and rotated callback tokens. Both are durable, but they are not one state machine. These are applications of the idempotent API principle: a retry carries stable or versioned identity, and the receiver distinguishes a replay from an authorized re-execution.[13][22]
Hosted fallback is deliberately asymmetric. Only job kinds with an explicit hosted policy may leave the desktop lane; at the present source snapshot, topic-bank refill is the only general worker job in the Fly-eligible registry. FAQ expansion and strategy reconciliation park for an eligible desktop rather than silently changing cost or trust boundaries. A local preference is only meaningful when the unavailable state is honest.
How the Strategy Adapts
“Self-adapting” here does not mean online weight training or an agent modifying its own code. It means that verified outcomes and fresher observations can produce a new, reviewable policy version. After report completion, bounded monthly or quarterly hooks can refresh inventory and baselines. Strategy reconciliation assembles nine deterministic strategy, evidence, and state blocks: current strategy, Search Console clusters, rank, competitor gap, PageSpeed, crawl/index evidence, the impact journal, monthly snapshot, and relevant algorithm updates. The job stores both the active strategy content hash and an input digest, coalescing duplicate reviews.[19]
Observation–memory–reflection–planning architectures have also been studied in simulated social agents, where ablations showed that each component contributed to believability.[26] That result motivates the memory-and-reflection shape, not its SEO correctness: our endpoint is governed production behavior, not believable simulation.
The result is not applied blindly. The finalizer validates a strict JSON shape and a closed finding vocabulary, confirms the active content hash and input digest, and rechecks source state. In Manual Review it stores the review and diff for a person; applying accepted amendments may then create an N+1 draft. In Full Autopilot it may create and activate N+1 through the normal lifecycle. An existing human-authored draft is never overwritten; it produces a deliberate no-op. If evidence changes during execution, the finalizer may enqueue exactly one successor rather than applying stale reasoning.
The same pattern drives narrower adaptations. A FAQ candidate requires recent ranked-keyword evidence and “People also ask” questions from the same provider run, a verified published URL, stored markdown, at least three clean unanswered questions, and a stable source hash. The agent returns plain answers to the verbatim questions; validation rechecks the markdown hash and emits one compound Page fix containing both visible FAQ content and corresponding FAQPage JSON-LD. It then travels through the shared fix and publication substrate.
FAQ is content utility, not a rich-result promise
Google stopped showing FAQ rich results in May 2026 and removed the feature documentation in June.[18] Visible FAQs may remain useful as reader-facing content when they answer demonstrated questions and improve page completeness. This architecture does not claim that FAQ schema itself produces a ranking or rich-result gain.
Finally, reports close the organizational loop. Generation freezes source, freshness, and comparison windows; QA determines whether the artifact is fit to publish; attachment—not generation—is the client-publication boundary. Only then does a concise strategy summary relay into the white-label portal. Client surfaces make no provider calls. Attached reports own report-derived metrics. The active-SEO card requires a deliberately published strategy summary, then derives bounded shipped, upcoming, and result rows from tenant/project-pinned impact-journal and active-backlog data. Map Pack is an independently gated projection.
Architecture Derived from Failure
The strongest design decisions did not originate in a whiteboard exercise. They came from adversarial audits and concrete failure paths in the August 2026 release. The corrections are useful because each one identifies a general class of agent-system error.[23]
Observation is not intent
Search Console queries had been able to masquerade as primary page intent. The repaired authority order is human > strategy > automatic observation; GSC refreshes target queries but cannot invent intent. This prevents a descriptive signal from becoming an unauthorized policy mutation.
Echo is not measurement
Daily rank-cache echoes could appear as fresh observations and contaminate monthly comparisons. The corrected snapshot policy distinguishes paid weekly measurement from cached presentation, lets measured evidence win on the same day, and normalizes location identity before comparison.
Retry is a financial event
Provider retries, partial fan-outs, ambiguous errors, and concurrent AI/DataForSEO lanes exposed undercount and double-count paths. A unified database ledger now serializes SEO budget admission under one tenant lock, preserves per-attempt identity, counts open reservations, settles idempotently, and fails closed on a blank budget.
A result is not finalized until recovery agrees
In-memory finalization could be lost after execution, RLS could turn an update into a zero-row “success,” and a machine-wide diagnostic could leak another tenant's activity. The durable outbox, selected-row verification, exact-token replay, runtime recovery, and sanitized heartbeat shape make completion and observability part of the security model.[22]
Identity must match material reality
A stale local-grid uniqueness constraint survived an earlier migration and blocked legitimate rescans; URLs differing only by advertising or attribution parameters fragmented PageSpeed and index-coverage identity. The repairs validate the actual database constraint and canonicalize observation identity before persistence rather than relying on a migration ledger or string equality alone.
These incidents show why an agent platform is a distributed system. The model may be the most visible component, but correctness depends on transactions, leases, identity, authorization, freshness, and recovery. Human-AI design guidance recommends making system status, limits, and correction paths legible to people; risk-management guidance similarly emphasizes traceability, measurement, and governance across the lifecycle.[16][17] In migrated paths those principles are enforced below the interface as data and state contracts, then surfaced to administrators as measured, stale, unavailable, queued, parked, failed, or completed states.
Why This Should Raise the Success Rate
Define a scored operational endpoint as the intersection of six contract conditions: the evidence is decision-eligible; the active strategy is the intended authority; the agent sees the right task projection; the action is authorized and contract-valid; the applicable publication or finalization side effect is idempotently receipted so a replay is distinguished from an authorized re-execution; and the outcome is verified and recoverably recorded.
We do not need to assume those failures are independent. This union bound applies only to the named contract-failure modes; it is not an upper bound on total agent failure, which may include omitted semantic, model, or environmental failures. Its value is decomposition: each named mode becomes observable and separately measurable.
- Epistemic control targets evidence error: where migrated, source, scope, time, method, and quality distinguish measurement from cache, echo, inference, absence, and staleness.
- Strategic control targets objective drift: one active version and an explicit authority order reduce opportunities for each agent or surface to invent its own goal.
- Attentional control targets context interference: for compiled-context lanes, a bounded task projection gives the model relevant evidence and acceptance criteria without asking it to discover the whole client state.
- Deliberative control targets unstructured reasoning error: staged research, typed outputs, evidence IDs, and closed vocabularies constrain semantic judgment while preserving it where rules are insufficient.
- Operational control targets distributed-side-effect error: compare-and-swap claims, attempt IDs, outboxes, callback tokens, deterministic Git receipts, and live verification provide bounded recovery paths toward one logical outcome.
- Governance control targets irreversible-policy error: manual diffs, protected human drafts, Full Autopilot gates, and tenant/client publication boundaries keep automation proportional to owner intent.
The hypothesized practical advantage over the prior pipeline is therefore not “more agents.” It is less ambiguity between them. Research produces a durable evidence graph; strategy resolves that graph into one operating policy; projections tell each lane what part it owns; the shared Page/Technical executor reduces divergent implementations; receipts give reporting something stronger than the agent's own assertion; and reconciliation changes future policy only after comparing the active hash with newly assembled blocks under their current per-block eligibility policies. Each boundary removes work the model previously had to infer.
Evaluation Protocol and Limitations
A credible next paper must measure the success-rate claim directly. AgentBench demonstrates why broad “model intelligence” scores are insufficient for interactive environments; our evaluation should use the real job contracts and score the entire trajectory, including recovery.[11] A credible protocol needs three linked studies: a paired semantic-quality replay for context compilation, an orchestration and fault-injection ablation for durability, and a longitudinal study of search outcomes. Without an explicit treatment date and prespecified counterfactual, that third study remains descriptive. No single comparison estimates the entire architecture.
- Unit of analysis: one typed job attempt against a frozen project, strategy hash, input digest, and acceptance contract.
- Primary endpoint: the correct terminal state for the job type—validated output or deliberate abstention, tenant/project binding, idempotent finalization, and, only for publication jobs, matching Git, deploy, and live proof.
- Secondary endpoints: unsupported-claim rate, stale-evidence use, context utilization, human edit distance, cannibalization/duplication violations, cost, wall time, recovery after injected faults, and live-verification rate.
- Context study: hold model, tools, project snapshot, and task constant; compare former versus compiled context across repeated samples under a fixed sampling configuration, recording model version, sampling parameters, and seed or system fingerprint where available; randomize order and blind human reviewers.
- Orchestration study: compare legacy and durable execution under component ablations and a fault matrix covering provider timeout, committed-but-lost response, stale hash, expired token, restart, zero-row update, duplicate delivery, and deployment failure.
- Longitudinal study: analyze comparable GSC, GA4, rank, grid, crawl, and, where configured and window-comparable, conversion outcomes after prespecified lags, separately from immediate task correctness; estimate causal effects only under a declared treatment and counterfactual design.
The study should report both task success and abstention quality. A system that refuses an unsupported FAQ, parks for a missing desktop capability, preserves a human draft, or marks an observation unavailable may complete fewer apparent actions while succeeding more often at the actual objective. “Did something” is therefore a poor endpoint; “reached the correct terminal state under the evidence and policy available” is the relevant one.
Code and evidence availability
The implementation trace was conducted against private Zyan source snapshot0249fd0321e4a383c2315eda6c4e6f56f0043d8d. Sources 19–23 name the relevant files, migrations, and contract families, but the source and tenant-scoped production receipts are not publicly inspectable. The founding chart and screenshot are public in Source 1; the screenshot's SHA-256 isa5491cb4910d5be7c4422d1b2a16311ff52ed9e67b1bd1edba8d82467a91715d. Exact daily Search Console rows, immutable raw export, and fixed-window boundary dates are not available in the public record, so the 17.1× calculation is internally reported rather than independently reproducible.
Current limitations remain substantial. We have no systematic agent-quality baseline; some freshness policies are not universal; one evidence table relies on artifact-level IDs rather than a direct run foreign key; legacy consumers sometimes collapse missing, stale, and read error; hosted fallback is not available for every job; Context Pack V2 depends on deployment configuration; and static source and tests are not fleet reliability. The founding 17.1× visibility observation is an association in a demand-volatile vertical, not a treatment effect. Those constraints narrow the conclusion but do not erase it.
Conclusion
The strategy compiler turns autonomous SEO from a sequence of model calls into a governed feedback system. Its hypothesized advantage is not that the agents know more in the abstract. The architecture is designed to tell agents—and, where consumer integration is complete, does tell them—which evidence is measured, cached, stale, or unavailable; which strategy is authoritative; which slice of the problem they own; which tools they may use; what constitutes completion; and how to recover when the world interrupts them. That is the difference between generating SEO work and operating a system that can adapt policy without discarding prior versions, human boundaries, or receipts.