Skip to content

Connection = Lock = Heartbeat

The defining architectural property of Aether is that the active gRPC stream connection simultaneously serves three roles: it is the distributed lock for an identity, the heartbeat proving that identity is alive, and the channel for all communication. When the stream closes — for any reason — all three are released atomically.

Distributed systems typically require three separate mechanisms to coordinate identity:

  1. A lock — to ensure only one process holds a given identity at a time (e.g., Redis SET NX)
  2. A heartbeat — to detect when a lock holder has crashed and release the lock
  3. A communication channel — for the process to actually exchange messages

Managing these separately creates surface area for bugs: stale locks when heartbeats stop, race conditions between lock expiry and reconnection, and operational complexity from three independent failure modes.

Aether collapses all three into the TCP/gRPC stream itself.

When a client opens a gRPC stream and sends InitConnection, the gateway:

  1. Authenticates the client
  2. Calls SET NX in Redis for the identity’s lock key with a 30-second TTL
  3. If the lock is already held → rejects with DuplicateIdentityError
  4. If the lock is acquired → the connection is the lock

The client does not call a separate lock API. There is no “acquire lock then connect” sequence. The connection attempt is the lock attempt.

While the stream is open, a background goroutine in the gateway refreshes the Redis lock TTL every 10 seconds. The TTL is 30 seconds, so a gateway crash gives a 30-second grace period before the lock auto-expires and the identity becomes available for reconnection.

Clients do not send heartbeat pings. The TCP keepalive and gRPC stream liveness are the heartbeat. If the client process crashes, the TCP connection closes, the gateway detects the EOF, and lock cleanup happens immediately — not after a TTL expiry.

When the stream closes (client disconnect, network failure, gateway shutdown, admin force-disconnect), the gateway:

  1. Calls DEL on the Redis lock key
  2. Decrements the workspace quota counter
  3. Updates any associated orchestration task state
  4. Writes the audit log entry

The lock is released synchronously as part of disconnect cleanup. Other gateway instances can immediately see the identity as available in Redis.

Two clients cannot hold the same identity simultaneously. The Redis SET NX operation is atomic. Even across a horizontally scaled cluster with many gateway instances, only one SET NX call wins. The second caller receives DuplicateIdentityError.

Principal types and their uniqueness constraints:

PrincipalUniqueness
AgentOne connection per workspace + implementation + specifier
Task (Unique)One connection per workspace + implementation + unique_specifier
Task (Non-Unique)Multiple connections allowed; server assigns unique IDs
UserOne connection per user_id + window_id (multiple tabs = multiple window IDs)
Workflow EngineOne active connection
Metrics BridgeOne active connection
OrchestratorOne connection per implementation + specifier
ServiceOne connection per implementation + specifier

When a client reconnects after a brief network partition, the previous lock may still be held in Redis (within the 30-second TTL). The client can include a resume_session_id in its InitConnection message to request atomic lock takeover.

The gateway executes a Lua script that atomically checks whether the existing lock belongs to the resuming session and replaces it if so. This prevents a race between the old lock expiring and a different client acquiring it.

The ConnectionAck response includes a resumed: true flag when this succeeds.

The TCP connection closes immediately. The gateway detects EOF, runs disconnect cleanup, and releases the lock. The identity is available for reconnection within milliseconds.

The TCP connection hangs. The gateway’s keepalive probe eventually times out and closes the stream. The lock TTL (30s) provides a window during which the client can reconnect with resume_session_id and atomically take over the lock without a gap.

The gateway process dies without running cleanup. Redis lock TTLs expire after 30 seconds. Clients reconnect to a different gateway instance. RabbitMQ Streams preserve message offsets so no messages are lost.

When restarting a container before the previous lock expires, clients receive DuplicateIdentityError. All three SDKs support a RetryOnDuplicate / retry_on_duplicate option that treats this error as recoverable and retries with backoff until the old lock expires (~30 seconds).

  • No separate heartbeat endpoint. There is no /api/heartbeat or keepalive message. If you are looking for one, you are working with a different system.
  • No zombie sessions. A session in the active sessions list is live by definition. Dead sessions expire within 30 seconds.
  • Lock visibility is binary. Either a client holds the identity or it does not. There is no “lock degraded” or “lock suspect” state.
  • Admin disconnect is immediate. The DELETE /api/connections/{session_id} endpoint sends a FORCE_DISCONNECT signal, which causes the stream to close and the lock to release synchronously.
ApproachLockHeartbeatChannel
AetherImplicit (stream = lock)Implicit (stream liveness)gRPC stream
Redis SET NX + heartbeat timerExplicit SET NXSeparate timer/cronSeparate
ZooKeeper ephemeral nodesEphemeral nodeSession keepaliveSeparate
etcd leasesLease grantLease keepaliveSeparate

Aether’s approach is simpler operationally because the failure modes for lock, heartbeat, and channel are identical — they all reduce to “is the TCP stream alive.”