Skip to content

Operational status codes for Honeycomb

Research date: 2026-10-02

Scope: Python libraries and current primary-source guidance for status telemetry covering Training Runs and recommendation serving.

Status: Research and implementation proposal; the example contract is not an approved specification or an implemented feature.

Recommendation

Extend the existing typed observations with a small, versioned diagnostic-code catalog using Python StrEnum and frozen dataclasses, then export its fields through the observability adapter. Keep the domain code, outcome class, retry decision, and alert policy separate. Use Honeycomb calculated fields to turn the observed outcomes into SLIs, and use freshness and presence signals to cover work that never finishes or never starts.

No reviewed library supplies a complete, generally applicable catalog for training and recommendation serving. There are real custom error-code libraries, but their documented contracts mostly concern API error responses. OpenTelemetry supplies transport and telemetry semantics; the service still owns the meanings of rec.training.published and rec.serving.fallback_used. This is a finding about the reviewed candidates, not proof that no other package exists.

The user clarified that training drill means a training process. This note uses the canonical term Training Run throughout; the two operation families are training and serving.

What the standards and platforms actually provide

OpenTelemetry span status has only Unset, Ok, and Error. Its tracing specification normally leaves successful instrumentation spans unset; explicitly setting Ok overrides error status. Custom application codes therefore belong in attributes, outside this enum. (Tracing API: Set Status)

The stable error.type attribute is intended to be predictable and bounded. OpenTelemetry explicitly recommends a domain-specific attribute when a domain has its own error identifiers, with error.type also set for failures. Successful operations omit error.type. (Error attributes)

OpenTelemetry's general error-recording guidance currently has Development status. It recommends consistent error classification on spans and metrics, a duration histogram covering both success and failure, and avoiding a failure classification for an outer operation that recovered through retry or graceful handling. A failed inner attempt can still have its own failed span. (Recording errors)

HTTP telemetry retains its literal transport meaning: http.response.status_code records the actual response code when one was sent or received. A successful HTTP response can carry a separate application outcome indicating reduced recommendation quality. (HTTP span conventions)

Honeycomb recommends dot-separated namespaces and stable field names for custom instrumentation. It recommends OpenTelemetry for Python instrumentation and also supports structured events through Libhoney. This means a custom status schema can use the supported telemetry path without a special Honeycomb status-code SDK. (Organizing data, Python options)

Related standards illustrate the distinction between lifecycle state and terminal result. OpenTelemetry's CI/CD conventions define pipeline state and result separately; the relevant attributes are Release Candidate and specifically describe software delivery pipelines. They are useful design precedents, but do not establish a general ML-training status vocabulary. (CI/CD attributes)

Google's approved AIP-193 is the closest established structured error-model precedent: canonical google.rpc.Status codes describe the broad failure, while ErrorInfo carries a stable (reason, domain) identity and structured metadata. Adopt that separation conceptually. The document governs Google-style APIs; it does not define successful/degraded recommendation outcomes or require this service to use protobuf. (AIP-193)

Python library assessment

The last column is this research's recommendation for this service, not a claim from the package author. Package availability and documented APIs were checked; packages were not installed or benchmarked.

Candidate Documented capability Fit for this change
Python enum.StrEnum and dataclasses String enum values; generated data-container methods; frozen=True rejects ordinary field assignment. StrEnum is available from Python 3.11. (Enum, dataclasses) Recommended catalog foundation. No extra dependency. Type annotations alone do not validate incoming JSON; validate at external boundaries.
googleapis-common-protos plus grpcio-status Python protobuf messages for Google's common APIs and conversion between rich google.rpc.Status and gRPC status. Python supports the richer protobuf error model. (Common protos, gRPC Python status API, gRPC error model) Closest established structured error model. Appropriate at a protobuf/gRPC boundary. Custom reasons go in ErrorInfo; canonical RPC codes remain canonical. Adding protobuf/gRPC solely for Honeycomb status fields is unnecessary.
opentelemetry-api, opentelemetry-sdk, opentelemetry-exporter-otlp Custom spans and attributes, counters and histograms, and OTLP export to Honeycomb. Python traces and metrics are stable; the Python status page still labels logs Development. (Python status, instrumentation, Honeycomb Python SDK) Recommended transport through the existing adapter. These packages do not define business status codes. Do not create another provider or exporter in domain modules.
Pydantic Validates enum members/values and serializes them to their underlying values in JSON mode. (Enum support) Useful if already used for public or persisted schemas. Optional for a catalog of internal constants; adding it only for this task is unnecessary.
APIException BaseExceptionCode supports named custom codes with messages/descriptions; FastAPI handlers map exceptions to response envelopes and logging. (Project source and README) A genuine custom-code helper, but its API-response/exception focus is narrower than the required outcome model. It does not supply training or serving degradation semantics. Avoid changing API responses or automatic logging merely to gain an enum wrapper.
returns Typed Result, Success, and Failure containers with composition helpers. (Result documentation) Useful if deliberately adopting result-oriented control flow. It neither defines codes nor supplies degraded/skip semantics or Honeycomb export. Too broad a control-flow change for instrumentation alone.
result / rustedpy/result A typed Rust-like result container; the repository now states it is no longer maintained. (Maintainer repository) Do not introduce it for this work.
structlog Structured dictionary-based logging with JSON and standard-library logging integration. (Official documentation) Useful for structured diagnostic records if consistent with the existing adapter. It supplies neither a catalog nor alert semantics.
prometheus-client Counter/histogram labels and an Enum instrument that represents one current state using one series per possible state. Enum instruments do not support its multiprocess mode. (Labels, Enum) Useful for an existing Prometheus pipeline. An Enum gauge is not a history of completed operations and cannot replace outcome counters/events. Prefer the existing OTel path here.
Tenacity Retries by exception or returned result, bounded stop conditions, backoff, and retry callbacks. (Official documentation) Reuse an existing retry mechanism if present. A retry library can report attempt outcomes, but does not determine whether a status is operationally actionable. Its bare decorator retries indefinitely, so explicit bounds are essential if adopted.
openlineage-python Standard lineage events and run states; its Python client documents one START event and one COMPLETE, ABORT, or FAIL event per run. (Python client reference) Useful when lineage interoperability is a separate requirement. These are workflow lifecycle states, not a recommendation-serving or alert catalog. Do not adopt a lineage subsystem just to label outcomes.
MLflow Its run status model has RUNNING, SCHEDULED, FINISHED, FAILED, and KILLED, with terminal-state classification. (RunStatus source) Reuse only if experiment tracking already requires it. A finished training run still needs a separate publication/quality outcome for this service.
Libhoney for Python Sends structured fields/events to Honeycomb's Events API. (Official Python guide) A valid transport alternative for an application already using it. Adding it alongside the existing OTel adapter would duplicate transport ownership without solving catalog design.

What this repository already implements

These are source observations from main at 991ec2a, not external recommendations:

Existing contract Evidence and implication
Typed operational results Observability seam defines OperationOutcome, FailureCategory, OperationName, and frozen event/metric observations. Preserve these fields and their consumers; the new catalog can add precision without replacing them.
Closed export schema Adapter maps outcomes to operation.outcome and failure.category, and explicitly allowlists attributes. Merely adding a Python field will not export it. Its error outcomes currently include only failure, unavailable, and lease_lost.
Fault-isolated runtime Runtime wraps the sibling telemetry implementation with FailSafeObservability. Export failure must remain unable to fail a request or a Training Run.
Worker observations Training worker records a terminal metric/event in _record_run_result for each claimed execution. Lease recovery can produce multiple executions of one logical run, so this count must not be assumed to be a deduplicated durable-run count.
Durable publication Storage, particularly _complete_publication_run, commits run success with snapshot publication. The worker performs quality assessment and retention afterwards. A later exception can yield worker-failure telemetry while the durable run remains succeeded; the new mapping must distinguish these facts.
Serving degradation API already reports insufficient evidence, stale snapshots, and personalization fallback. A stale response can still return HTTP 200. Staleness is recorded as a separate degradation observation rather than necessarily changing the operation's original outcome.
Future diagnostic contract Domain glossary defines Diagnostic Occurrence; prototype explores the richer code/effect/retry contract. This is not yet the complete production diagnostic implementation. Reuse its vocabulary without claiming the prototype is shipped.
Existing alert intent Operations contract already describes serving errors/latency, worker readiness, lease loss, training failure, publication stalls, and drift. Runbooks keep scope and entity identifiers out of metric labels. Research examples do not amend these policies.
Existing dependencies and tests Package configuration already includes Pydantic and the sibling telemetry package, with an OpenTelemetry extra. Adapter contracts use in-memory span, log, and metric exporters; unit observations and operations tests provide the nearby verification seams.

The sibling telemetry source at ../telemetry/src/telemetry/otel.py starts spans with context captured on entry and exposes set_outcome(outcome, failure_category) for their final result. Its structured events travel through the logging path. Rebinding context after computing a result does not retroactively add that result to the existing span. It currently sets successful operation spans to Ok and does not automatically derive error.type from failure.category. Any alignment with newer error-recording guidance is a separately tested runtime change.

Proposed status contract

Use explicit, stable string values, for example rec.training.source_timeout, in app.recommendations.status_code. A human-readable code is already machine searchable. If compact numbers are desired later, add a permanently assigned numeric alias from the same catalog; do not derive alert meaning from arithmetic ranges or reuse HTTP values in HTTP fields.

Separate these dimensions:

Dimension Proposed meaning and bounded values
Operation training or serving.
Measurement kind request, run, or attempt; the denominator of each query must select one.
Outcome class success, degraded, failure, skip. All are terminal observations.
Status code A stable catalog entry identifying what happened. Never use exception messages or dynamically generated identifiers as codes.
Reason code Optional bounded cause, such as personalization_unavailable or source_timeout, when several causes share an outcome code. Avoid an independent reason field when it merely duplicates the code.
Retryability Whether repeating this operation could help, under its contract. Keep a separate retry_scheduled decision or attempt count when relevant; exhaustion, cancellation, budgets, and idempotency can prevent a retry.
Lifecycle state Existing workflow state such as queued/running/publishing. It is not a terminal outcome or a measure of serving health.
Action severity An alert-policy result such as observe/investigate/page, based on sustained impact, freshness, scope, environment, and runbook. A retryable error need not page; a nonretryable corruption may require immediate action.

degraded means usable output was produced under a reduced-quality or fallback path. It is not automatically good or bad for every SLI. skip means the operation was intentionally not performed; it is not evidence of fresh output. Expected cancellation may be represented as a documented skip reason; missing a required deadline remains a failure of the relevant freshness objective.

Publication and model quality are separate facts. Drift or an evaluation warning can accompany a successfully published snapshot; report the measured quality condition and threshold separately. Likewise, retention or cleanup failing after publication must not emit publication_failed or erase publication success. Record its diagnostic condition under its own operation, without adding another Training Run completion to the SLI denominator.

These entries are examples to reconcile with existing behavior and product policy. The no_new_evidence example is hypothetical; it does not authorize a new skip path. A worker's NO_WORK polling result is a scheduler observation, not a skipped Training Run:

Example code Class Evidence needed before recording it
rec.training.published success Complete Recommendation Snapshot publication committed successfully. Completing fitting alone is insufficient.
rec.training.no_new_evidence skip The scheduled operation intentionally skipped under a documented no-change rule. It must not advance last-successful-publication time.
rec.training.source_timeout failure The measured attempt or final run ended because source access timed out; distinguish those measurement kinds.
rec.training.publication_failed failure Publication did not complete; the previously published serving head remains the relevant head.
rec.serving.snapshot_served success The serving contract was satisfied from the published snapshot/projections.
rec.serving.fallback_used degraded An allowed fallback produced usable output; attach its finite reason and evaluate quality separately.
rec.serving.stale_snapshot_served degraded A stale snapshot was deliberately served under the allowed serving policy; HTTP 200 does not establish freshness.
rec.serving.snapshot_unavailable failure No usable published snapshot was available to fulfill the request.

Freeze wire values once released and use explicit literals rather than order-dependent values. Keep schema_version for event shape and a catalog version for definitions. A changed meaning requires a new code or an explicit version migration. Provide an unknown/unclassified fallback at the observability boundary so unexpected errors remain visible without exporting their raw messages. Application logic must not silently accept unknown codes as successes.

This small Python sketch illustrates catalog metadata only; it is not a proposed replacement for the repository's existing OperationOutcome or an importable production API:

from dataclasses import dataclass
from enum import StrEnum
from typing import Literal

class DiagnosticCode(StrEnum):
    SERVING_FALLBACK_USED = "rec.serving.fallback_used"
    TRAINING_SOURCE_TIMEOUT = "rec.training.source_timeout"

@dataclass(frozen=True, slots=True)
class CodeDefinition:
    code: DiagnosticCode
    status_class: Literal["success", "degraded", "failure", "skip"]
    retryable: bool
    runbook_key: str

SOURCE_TIMEOUT = CodeDefinition(
    DiagnosticCode.TRAINING_SOURCE_TIMEOUT, "failure", True, "training-source-timeout"
)

The minimum additive fields are schema version, diagnostic code, and any required bounded reason; reuse the existing operation, outcome, failure category, and scope context. Generate class and retryability metadata consistently from the catalog rather than letting callers contradict it.

Proposed flat event schema

This illustrative completed-request event represents a permitted fallback with HTTP 200. Every field below is either a scalar or a bounded code. Scope labels are illustrative, configuration-owned labels for the complete Data Source / Tracking ID / Catalog ID tuple; production naming must match the repository's approved observability contract.

{
  "service.name": "commerce-recommendations-api",
  "deployment.environment.name": "production",
  "app.recommendations.schema_version": 1,
  "app.recommendations.catalog_version": 1,
  "app.recommendations.event": "operation.completed",
  "app.recommendations.operation": "serving",
  "operation.outcome": "personalization_fallback",
  "app.recommendations.measurement_kind": "request",
  "app.recommendations.status_class": "degraded",
  "app.recommendations.status_code": "rec.serving.fallback_used",
  "app.recommendations.reason_code": "personalization_unavailable",
  "app.recommendations.retryable": false,
  "commerce.data_source.id": "source_example",
  "commerce.tracking.id": "property_example",
  "commerce.catalog.id": "catalog_example",
  "app.recommendations.duration_seconds": 0.018,
  "app.recommendations.serving.response_usable": true,
  "app.recommendations.serving.freshness_satisfied": true,
  "app.recommendations.serving.snapshot_age_seconds": 900,
  "http.response.status_code": 200
}

For this example, leave the domain span status unset and omit error.type because the operation fulfilled its availability contract through a permitted fallback. A failed personalization suboperation can have its own error status. If fallback violates the defined operation contract, classify that operation as failure instead. For a final Training Run timeout, set its domain span to Error and add a bounded error.type, such as rec.training.source_timeout, consistently with the failure metric. Preserve existing HTTP instrumentation semantics on the HTTP span.

This is a proposed future OTel mapping. Structured completion events and terminal span attributes are different export paths: if the generic telemetry API cannot enrich a span at completion, extend that API and its adapter explicitly. Setting attributes directly in workers would bypass the required adapter boundary. Preserve legacy outcome fields during any additive schema rollout.

Only service-generated, approved run/attempt/trace identifiers may be added for diagnostic correlation. Keep them out of metric labels. Do not export Shopper identifiers, raw Order IDs, raw Browsing Session IDs, source rows, SQL values, credentials, bearer material, or exception messages containing them. Scope labels must come from trusted configuration/context, never an unchecked request field. A redaction hook is not a substitute for this allowlist.

Proposed metric schema and counting rules

Metric Instrument / unit Measurement and labels
recommendations.operations.completed Counter / {operation} Increment once at the chosen terminal boundary. Labels: bounded operation, measurement kind, class, and catalog code. No scope or individual entity identifiers.
recommendations.operation.duration Histogram / s Record once with the same outcome classification; add error.type only for failures. Include successes and failures in one instrument.
recommendations.snapshot.unavailable Observable gauge / {scope} Fleet count of expected scopes without a usable snapshot, without identity labels. Correlated scoped events identify affected scopes.
recommendations.snapshot.stale Observable gauge / {scope} Fleet count of overdue published snapshots. Failed/skipped training cannot reset publication age. Correlated scoped events carry actual ages.
recommendations.worker.heartbeat_timestamp Observable gauge / s Last verified heartbeat/progress for a bounded process role, emitted by a surviving observer where possible; no per-job/run identifiers.

This schema is a proposal. Reuse existing equivalent metrics rather than emitting duplicates. Counters are convenient for rates; histogram counts can also supply throughput when the metric backend exposes them. Metric measurements are made independently of span sampling.

The repository forbids Data Source, Tracking ID, Catalog ID, run, snapshot, product, and user identifiers in metric labels, including hashed representations. Keep the complete Commerce Scope in approved event/trace context. A bounded observer should inspect each expected scope's control metadata and emit scoped age, availability, and last-success observations; fleet metrics above are supporting summaries, not a replacement for per-scope detection. Put a hard bound on the remaining metric-label combinations. Prometheus's instrumentation guidance specifically cautions that each label combination adds a time series. (Instrumentation and cardinality)

Process presence alone is insufficient: inspect pending-work age and owned-run progress/lease state as well, so one healthy worker cannot conceal an overdue run or a scope that never publishes.

Define one owner for the terminal observation: request completion for serving, committed terminal state for a Training Run. Failed inner attempts produce attempt observations; only the final logical run contributes to the run-success denominator. Never count every span, start event, heartbeat, retry, and completion together as operations.

“Exactly once” here means one measurement call per terminal boundary in the application, not a guarantee of exactly-once delivery to Honeycomb. A process can die before emitting, exporters can lose or retry data, and a retrying worker can revisit completion logic. Use a guarded terminal transition and test duplicate callbacks. If audited completeness is required later, derive it from durable run state with a bounded reconciliation/outbox design and explicit duplicate handling; that is a separate reliability requirement, not something an enum or OTel exporter provides.

Periodic freshness and heartbeat coverage are essential because a missing completion event does not prove success. Prometheus's primary guidance identifies last success as a key batch-job signal and recommends progress/heartbeat evidence for offline processing. (Batch and offline processing guidance)

Honeycomb queries and later alerting

Honeycomb SLIs are calculated fields returning true for good qualified events, false for bad qualified events, and null for events outside the population. The two-argument IF form supports this qualification directly. (Create an SLI)

A proposed serving-availability SLI is below. It assumes the production dataset/environment has already been selected, and every qualified completion has the required fields. The denominator includes permitted fallbacks; a separate quality SLI can require the appropriate quality outcome. Each scope-specific SLO must also qualify on the complete approved Commerce Scope.

IF(
  AND(
    EQUALS($app.recommendations.event, "operation.completed"),
    EQUALS($app.recommendations.operation, "serving"),
    EQUALS($app.recommendations.measurement_kind, "request")
  ),
  EQUALS($app.recommendations.serving.response_usable, true)
)

Validate missing-field behavior with Honeycomb's preview before deployment. A missing outcome must not silently remove a failed request from the denominator. Observe schema-invalid/unknown outcomes separately. Set a latency criterion only after selecting the promised latency budget; do not invent a target from an HTTP status analogy.

The proposed alert families are:

Objective Suggested evidence and policy
Serving availability and latency SLO burn over qualified request outcomes and latency; distinguish valid client rejections from service failure under an explicit policy.
Serving quality/fallback Dedicated quality SLI or sustained fallback proportion, using a valid-request denominator and minimum traffic appropriate to the deployment. An isolated acceptable fallback need not page.
Training health Final failed-run observations for diagnosis; publication age/deadline per expected scope for impact. A run failure with a still-fresh serving head may warrant investigation before paging.
Silent training or worker outage Independent heartbeat/presence and overdue-publication signals, including scopes that have never published. A worker cannot report its own death reliably.

Honeycomb offers query-threshold Triggers, including configurable evaluation windows and consecutive threshold breaches. Its current metrics Trigger documentation supports temporal aggregation through calculated fields such as RATE() or INCREASE(); these differ from ordinary event-count queries. (Create a Trigger, Metrics Triggers)

For serving objectives, alert on consumed error budget and sustained user impact. Google SRE's primary guidance develops multiple burn rates and short/long windows to balance detection and reset time. Honeycomb has its own Budget Rate and Exhaustion Time burn alerts; configure their documented semantics rather than assuming a Prometheus rule is copied verbatim. (Google SRE: alerting on SLOs, Honeycomb burn alerts)

Sampling matters: Honeycomb weights supported aggregations using SampleRate; its COUNT_DISTINCT does not correct for sampling. Weighted estimates do not reconstruct a rare failed run that was never sent. Prefer retaining all low-volume Training Run completions and freshness/presence observations. For high-volume serving, keep accurate unsampled metrics and ensure any sampled event SLI preserves the correct sample-rate metadata and denominator. (Sampled data in Honeycomb)

A head-sampling decision occurs before the final outcome is known, so setting an error attribute at completion cannot recover an already discarded trace. Tail sampling can use final attributes, but inherits head-sampling losses and must handle delayed/long-running training spans within bounded buffers. Keep alert-critical training completion/freshness observations on an explicit retention path rather than relying on an arbitrary long trace being retained. (Honeycomb sampling, Tracing API and sampling timing)

Before enabling alerts, verify no-data behavior in the actual Honeycomb plan and query shape. The reviewed Trigger pages did not establish a universal rule for absent groups. Do not assume a grouped COUNT < 1 query will materialize every missing scope. A bounded observer of expected scopes should emit explicit stale/unavailable evidence, and an independent presence check should detect loss of that observer or its export path.

Implementation path for this repository

The smallest first release is an additive status catalog and structured completion events through the existing log export path. It can be queried in Honeycomb without first changing the shared runtime's span interface. Confirm that the deployment enables and routes the log signal: these application events are not automatically attributes on every trace span.

  1. Define explicit diagnostic codes and catalog metadata beside the application observability seam, using the existing outcomes and classified failures as inputs. Keep the catalog free of SDK imports. Resolve fallback/rejection/skip semantics before freezing their wire meanings; do not add new product behavior solely to fill a classification table. Model severity, effect, and retry conditions separately if the richer Diagnostic Occurrence contract is adopted.
  2. Add typed optional code fields to EventObservation or a dedicated typed completion observation. Extend EventName only where no existing event represents the measurement. Update _EVENT_ATTRIBUTES, _event_attributes, and the observation translation in the adapter. Preserve the existing operation.outcome, failure.category, and commerce.data_source.id / commerce.tracking.id / commerce.catalog.id context keys. The app.recommendations.* names above are proposed new fields, not fields already exported.
  3. Give each measurement exactly one owner. Enrich existing worker-attempt completion events; derive logical-run success from the committed run/publication transition. For serving, collect facts from anchored, global, personalized, and swimlane paths into a single request completion result, including error and admission paths. Select one event family for each SLI denominator so legacy failure events and new completion events are not both counted as requests.
  4. Define deterministic classification for coexisting conditions. For example, one response can use a fallback and a stale snapshot: choose one primary code by documented precedence while retaining separate bounded freshness/quality facts. Preserve publication success across subsequent maintenance failure. A failed worker attempt and a recovered successful run must remain distinguishable.
  5. Reuse equivalent existing metrics. Add only reviewed finite dimensions for codes/classes; never put scope, run, snapshot, product, or Shopper identifiers into metric labels. If codes are required on completed spans as well, extend the sibling runtime with a schema-validated, backend-neutral final-attribute interface and adapt it here. Do not import OTel directly in workers or serving code, open a second provider, or use a late context bind as a substitute.
  6. Version the changed telemetry contract in schema versions. Retain old fields through consumer migration. A telemetry-only addition need not change database or public API schemas; persisting full Diagnostic Occurrences or exposing them through APIs requires their own contract/version decision. Existing schemas are all version 1 at the inspected revision.
  7. Before alert activation, add the bounded observer for expected-scope publication freshness and pending/stalled work, using control metadata. Define run/attempt identity and duplicate handling if durable event delivery is required. Validate representative events and SLI populations in Honeycomb, then amend the operations contract/runbooks with agreed policy and recovery rules.

No new status-code dependency is needed for this first release. Public HTTP diagnostics, if later requested, can project the same catalog through an appropriate API error format; that need not change the status fields used for monitoring.

Implementation evidence to add later

Boundary Observable evidence required
Catalog and adapter Exact stable wire values, bounded attributes, rejection of sensitive/unregistered fields, compatibility of existing fields, and exporter failure isolation.
Serving One completion per request across success, failure, rejection, insufficient evidence, stale output, and personalization fallback; HTTP code remains the actual transport result.
Training Attempt versus logical-run counts, successful publication, failed publication preserving the prior head, lease recovery, and post-publication cleanup failure without a false publication-failure code.
Missing work Never-published and overdue scopes, stopped workers, interrupted runs, missing exporter data, and loss of the freshness observer remain detectable.

Use the existing in-memory exporters and domain fakes for these checks. Run the focused test first, then make test for the changed production boundary. Run make migration-check only if durable storage changes. This research does not claim those proposed implementation tests have passed.

Research verification

  • make docs-check: passed, 73 Markdown files checked.
  • make test-focused TEST=tests/delivery/test_repository_harness.py::test_agent_entry_points_and_verification_assets_exist: passed, one test. This verifies the documentation entry-point assets, not proposed runtime behavior.
  • Graph navigation used graphify query, graphify reflect --if-stale, and graphify save-result; all completed. Their generated metadata is separate from production code.
  • git diff --check and the new-file git diff --no-index --check reported no whitespace errors.
  • No production, migration, load, or live Honeycomb checks were run because this change is research documentation. No Python dependency was installed and no alert was activated.

Search scope and limitations

The search covered official OpenTelemetry specifications/Python guides, Honeycomb documentation, Python standard-library documentation, maintainer repositories/docs for the tabled libraries, OpenLineage, MLflow, and Google SRE. Searches included Python custom status/error-code libraries, typed result catalogs, OTel error semantics, Honeycomb custom attributes/SLIs/Triggers/sampling, and run lifecycle standards. Search results led to primary documentation; third-party comparisons were not used as evidence.

The finding is strongest for contract design and documented capabilities. It does not certify package maintenance, security, compatibility with the repository's lockfile, benchmark overhead, or Honeycomb plan entitlement. Documentation can change after the research date. The live Honeycomb account, collector configuration, and deployed sampling were not verified. No dependency installation, application code change, hosted configuration change, or external telemetry submission was performed for this note.