MESH ONLINECODENAME: Paranoid
v0.36
reference

Error Codes

This page enumerates every error type the core crate surfaces. Errors are organized by the operation they come from — ingestion, consumption, and adapter — and each variant includes the conditions under which it fires and the right response from a caller.

The crate uses thiserror throughout, so every variant has a Display impl and (where it wraps another error) a working source() chain. Pattern match on the variant when you need to make a decision; format the Display when you need to log.

IngestionError

Returned from EventBus::ingest() and EventBus::ingest_raw().

VariantDisplayWhen it firesWhat to do
Backpressure"backpressure: ring buffer full"Shard's ring buffer full + backpressure policy rejected the eventApply your retry policy; the bus will not surface this if Block mode
Sampled"event dropped due to sampling"Sampling/decimation policy dropped the event before it reached a shardExpected under sampling; no caller action needed
Unrouted"event has no routable shard"Hashed shard id is not in the routing table (e.g. mid-scaling)Back off briefly and retry — topology stabilizes within milliseconds
ShuttingDown"event bus is shutting down"Bus is in shutdown; new ingests rejectedStop ingesting; flush downstream state and exit
Serialization(_)"serialization error: ..."Event payload couldn't be serializedBug — investigate the payload; the error's source chain points at the underlying serde_json::Error

Unrouted is distinct from Backpressure so callers can apply the right remediation. Backpressure says "the destination is full"; unrouted says "there's no destination right now." Pre-fix versions of the bus collapsed these into one variant, and callers applied back-off-and-retry to unrouted errors that wouldn't be fixed by waiting — they needed to retry until the topology settled, which is a different shape of retry.

ConsumerError

Returned from EventBus::poll().

VariantDisplayWhen it firesWhat to do
Adapter(_)"adapter error: ..."Underlying adapter failed; the wrapped error is the adapter'sSee AdapterError below; is_retryable() says whether to retry
InvalidCursor(_)"invalid cursor: ..."Cursor in the request couldn't be decodedDon't pass that cursor again; start from current tail with no cursor
InvalidFilter(_)"invalid filter: ..."Filter in the request couldn't be parsed or evaluatedBug — investigate the filter; the message includes a parse position

A ConsumerError::Adapter wraps an AdapterError, so the full classification surface is available through the wrapped error. Use From<AdapterError> to convert, or pattern match on the wrapper.

AdapterError

Returned from adapter operations (Adapter::on_batch, Adapter::poll_shard, Adapter::flush, Adapter::shutdown). Also wrapped in ConsumerError.

VariantDisplayWhen it firesClassification
Transient(_)"transient error: ..."Retryable failure (timeout, transient network issue)is_retryable() == true
Fatal(_)"fatal error: ..."Unrecoverable stateis_fatal() == true
Backpressure"backend backpressure"Backend rejected for capacity reasons (Redis MAXLEN, JetStream MaxBytes, etc.)is_retryable() == true
Connection(_)"connection error: ..."Connection-level failure (refused, broken, reset)Not retryable by default — covers both transient ("send failed") and permanent ("not initialized") cases without distinguishing
Shutdown"adapter is shut down"Adapter was asked to stop and is no longer accepting workis_shutdown() == true; distinct from Connection so callers can tell "we asked it to stop" from "transport failure"
Serialization(_)"serialization error: ..."Adapter couldn't serialize/deserialize event dataNot retryable; bug in payload or adapter codec

Classification methods

rust
impl AdapterError {
    pub fn is_retryable(&self) -> bool;
    pub fn is_fatal(&self) -> bool;
    pub fn is_shutdown(&self) -> bool;
}

The bus's dispatch loop reads these to decide what to do with a failed batch:

  • Retryable. The batch is requeued with an exponential backoff up to a bounded number of attempts.
  • Fatal. The batch is dropped, the bus's stats record the drop, and the error is logged at error level.
  • Shutdown. The batch is dropped and ingestion is halted; the bus's shutdown is presumed to be in flight.
  • Connection (default). Conservatively non-retryable. The bus skips the retry loop and drops the batch immediately. This avoids burning the retry budget on a backend that's gone for good.

The default decision for Connection errors is conservative on purpose. If you know your backend's connection errors are transient and you want them retried, return AdapterError::Transient(...) from your adapter instead.

Subsystem-specific errors

Beyond the core trio, individual subsystems define their own error types. The ones most likely to surface in application code:

ScalingError

The shard-mapper scaling error. The EventBus scaling methods are manual_scale_up(count: u16) / manual_scale_down(count: u16), and both return AdapterError; ScalingError is the mapper-level type underneath.

VariantWhen it fires
InvalidPolicy(_)The scaling policy was rejected — the string says why
AtMaxShardsAlready at the configured shard ceiling
AtMinShardsAlready at the shard floor
InCooldownA scale operation is still inside its cooldown window; retry later
ShardCreationFailed(_)The new shard couldn't be built — investigate

ConfigError

Returned from EventBusConfigBuilder::build().

One variant — match it, read the string to see which knob.

VariantWhen it fires
InvalidValue(_)A configuration value was rejected; the string names the offending setting — a num_shards count out of range, inconsistent batch sizing (max_events == 0), a feature requested but not compiled in, …

Adapter-specific errors

Each shipped adapter has its own error type with backend-specific variants. The most useful ones from each:

  • NetAdapterNetAdapterError::SessionFailed, NetAdapterError::RoutingFailed, NetAdapterError::AuthRejected.
  • RedisAdapter — wraps redis::RedisError with classification ("retryable" for read timeouts and replica failovers, "fatal" for auth failures).
  • JetStreamAdapter — wraps async-nats::Error with similar classification.

These are surfaced through AdapterError::Connection, AdapterError::Transient, or AdapterError::Fatal as appropriate, so callers don't need to know the specific backend to apply the right policy. Match on the AdapterError variant, not on the inner error, unless you have a backend-specific reason.

TokenError

Returned from the channel-auth token issuance and verification paths in net::adapter::net::identity.

VariantWhen it firesWhat to do
InvalidSignatureThe token's signature doesn't verifyReject; the credential is forged or corrupted
InvalidFormatWire bytes are too short or malformedReject; the credential is corrupted or garbage
ExpiredToken's not_after is in the past, modulo the configured clock-skew windowRe-issue from the current holder; tokens are time-bound on purpose
NotYetValidToken's not_before is in the futureWait, or re-issue with an earlier validity window
NotAuthorizedNo valid token covers the requested actionRequest a token with the right scope (publish, subscribe, admin, delegate)
DelegationNotAllowedThe token lacks the DELEGATE scope but tried to re-delegateIssue from a token that carries delegate authority
DelegationExhaustedDelegation depth hit zero and the token is being re-delegatedThe chain has run out of remaining delegation hops
RevokedA chain link is at or below its issuer's revocation floorRe-issue; kept distinct from NotAuthorized so you can tell a revoked credential from a never-authorized one
ReadOnlySigning was attempted with a public-only (zeroized / read-only) keypairYou hold a verify-only key — it can't sign
ZeroTtlduration_secs == 0 was passed to try_issueIssue with a non-zero TTL (a 0-TTL token is instantly Expired)
TtlTooLongRequested TTL exceeds the ceiling (MAX_TOKEN_TTL_SECS)Issue inside the bound; issue_token soft-clamps, try_issue returns this so callers can decide

The TTL ceiling is a hard cap on the auth surface — issuing a token past one year is rejected on the fallible path and clamped on the SDK's infallible path. Long-lived grants need periodic re-issue, which re-checks the issuer's signing key and current policy.

TagMatcherError

Returned from capability-tag matchers when the requested matcher can't be compiled or evaluated.

VariantWhen it firesWhat to do
RegexNotBuiltIn { pattern }A TagMatcher::Regex was used against a build compiled without --features regex; carries the offending patternRebuild with --features regex or use a non-regex matcher kind

The regex Cargo feature is off by default — regex matching adds about 1.1 MiB to binding artifacts, and most callers don't need it. Builds that do can opt in. Pre-v0.24 the regex-less fallback silently returned empty matches, which made misconfigured queries look indistinguishable from "no entries match"; v0.24 replaced that with the structured error above.

nRPC errors (RpcError)

Returned from call_typed, call_streaming_typed, call_client_stream_typed, call_duplex_typed, and the underlying MeshRpc surface.

VariantWhen it fires
RpcError::NoRoute { target, reason }The target node id is unknown to the local mesh, or the reply-channel subscription couldn't be set up
RpcError::Timeout { elapsed_ms }The call's deadline elapsed before a response arrived (the caller emits CANCEL)
RpcError::ServerError { status, message, headers }The server returned a non-Ok RpcStatus; headers carries the structured sidecar (e.g. a net-failure-schematic verdict)
RpcError::Transport(_)Underlying transport error (publish failure, encryption, …); wraps AdapterError
RpcError::Codec { direction, message }Request/response failed to encode/decode; direction is CodecDirection::{Encode, Decode}
RpcError::CapabilityDenied { target, capability }The callee refused because the caller lacked the required capability
RpcError::CancelledA MeshNode::cancel(token) aborted the in-flight call

There is no NoServer / NoMatchingServer / Panic variant — a handler that panics or returns a typed application error surfaces as ServerError { status, message, headers }, carrying the wire-stable status codes NRPC_TYPED_BAD_REQUEST / NRPC_TYPED_HANDLER_ERROR (part of the cross-language fixture). The binding-native typed wrappers (TS / Python / Go) re-raise these as idiomatic exceptions.

Organization errors (the org: vocabulary)

Returned from mesh.org(..) and OrgClient::call — see Private capabilities. Unlike every other error on this page, this one is a string vocabulary rather than a Rust enum, because it has to survive four FFI boundaries unchanged. It is single-sourced from OrgSdkError::to_wire, pinned by net/crates/net/tests/cross_lang_org/error_vectors.json, and re-parsed identically by the Node, Python, Go, and C bindings.

The shape is org:<domain>:<kind>[: <detail>]. The detail is human-facing and must not be parsed for semantics.

DomainLocal?Meaning
credentialsyesThe local credential set could not authorize this call. Nothing was sent
discoveryyesNo provider this credential set is authorized to call was found. Nothing was sent
admission_deniednoA provider's admission engine evaluated the call and refused it
rpcnoTransport, or a server error that is not an admission denial
unknownnoParser / ABI fallback — this binding's vocabulary disagrees with the build

The domain is the load-bearing fact, and is_local is the question it exists to answer: did anything leave this process? Every binding exposes it directly — domain.is_local() in Rust, ParsedOrgError.is_local in Python, OrgError.IsLocal() in Go, and distinct negative return codes in C — so you never have to re-parse the message to find out.

Two properties of this vocabulary are deliberate and worth understanding before you write code against it.

A binding that cannot classify a string reports unknown, never one of the four canonical domains. Reporting admission_denied for an unparsed string would assert that a request reached a provider and that provider's admission engine ran — a claim the binding is in no position to make. unknown appearing in your logs means a version skew between a binding and the build it is talking to, not an authorization problem.

Remote denials are coarse on purpose. admission_denied carries exactly one of denied, not_supported, or unavailable, with no detail at all. A precise remote reason would be a credential oracle — an attacker could walk it to learn which part of a credential set was wrong. The detailed reason is recorded provider-side for audit and never crosses the wire. Do not write caller logic that branches on a finer remote reason; there isn't one.

The org:rpc: kinds reuse the RpcError vocabulary above (timeout, no_route, cancelled, server_error, transport, codec_encode, codec_decode, capability_denied) rather than minting second names for the same conditions.

Subnet authority errors (the subnet: vocabulary)

Returned from subnet-authority configuration and provisioning — constructing a mesh with subnet_authorities / subnet_exports, installing gateway credential sets, declaring boundaries, applying control facts, and resolving a named export at serve time. Like the org: vocabulary it is a string envelope rather than a Rust enum, single-sourced from Rust, pinned by net/crates/net/tests/cross_lang_subnet/stable_kinds.json, and consumed by the Node, Python, Go, and C bindings.

The shape is subnet:<kind>[: <detail>], and unlike org: there are no domains — because every subnet: error is local and startup-shaped: a configuration, decode, or install was refused before (or without) any node-state mutation. Nothing ever left the process. A remote refusal of an exported call is never a subnet: error; it surfaces through the org: taxonomy above, so a caller cannot probe a provider's subnet configuration through error kinds.

Two bands of kinds:

BandExamplesSource
Core verifier refusalsunknown_authority, wrong_topology_epoch, scope_not_ancestor, revoked, issuer_attenuation_broadened, expired, invalid_signature, invalid_formatthe credential/fact verifier — why an artifact was refused
Local configurationunknown_export_name, duplicate_export_name, empty_authority_roots, invalid_id_hex, path_too_deep, invalid_path_level, invalid_accessthe binding conversion layer — why a configuration never reached the core

Three properties worth knowing before writing code against it:

  • Serve-registration failures wrap the envelope (… failed: subnet:unknown_export_name: …) rather than leading with it. Every binding classifier SCANS for the token rather than requiring it at position 0, and the token ends at the next colon or whitespace: parseSubnetKind / classifySubnetError (Node), net.subnet.parse_subnet_kind or the raised exception's .kind (Python), ParseSubnetKind / errors.Is(err, ErrSubnet) (Go), the NET_ORG_ERR_SUBNET return code plus the out_err wire (C). Classify on the type and kind — never on message text, which carries operator prose that is not part of the contract.
  • An unrecognized kind passes through verbatim as data. A binding never remaps a kind it does not know onto one it does — the same anti-counterfeiting rule as org:'s unknown domain.
  • applied: false from a control fact is not an error at all. It is an authenticated stale/idempotent outcome — the fact verified but changed nothing. Don't retry it.

Per-peer stream errors (StreamError)

Returned from the per-peer stream API on MeshNode.

VariantWhen it fires
StreamError::BackpressureThe stream's outbound queue is full — no packets were enqueued; retry, drop, or surface. The retry-safe case (send_with_retry handles it automatically)
StreamError::NotConnectedThe underlying session is gone (peer disconnected, never connected, or the stream was closed)
StreamError::Transport(_)Underlying transport failure (socket / encryption error); wraps the adapter-level error's message

WindowFull and stream reset live below this surface — the tx-credit admittance value and the SUBPROTOCOL_STREAM_RESET wire message, respectively — they are not StreamError variants.

A note on credentials in URLs

Adapter constructors and Debug impls scrub user:password@ from connection URLs before logging or rendering. A misconfigured operator who put a password directly in the URL won't leak it into log sinks — the redactor identifies the rightmost @ in the authority component and replaces the userinfo with [REDACTED].

This is per-adapter behavior, not part of the error API itself, but it shows up in Debug output of every adapter config and is worth knowing about when reading logs.