Connection = Lock = Heartbeat
The defining architectural property of Aether is that the active gRPC stream connection simultaneously serves three roles: it is the distributed lock for an identity, the heartbeat proving that identity is alive, and the channel for all communication. When the stream closes — for any reason — all three are released atomically.
The Problem This Solves
Section titled “The Problem This Solves”Distributed systems typically require three separate mechanisms to coordinate identity:
- A lock — to ensure only one process holds a given identity at a time (e.g., Redis
SET NX) - A heartbeat — to detect when a lock holder has crashed and release the lock
- A communication channel — for the process to actually exchange messages
Managing these separately creates surface area for bugs: stale locks when heartbeats stop, race conditions between lock expiry and reconnection, and operational complexity from three independent failure modes.
Aether collapses all three into the TCP/gRPC stream itself.
How It Works
Section titled “How It Works”Lock Acquisition on Connect
Section titled “Lock Acquisition on Connect”When a client opens a gRPC stream and sends InitConnection, the gateway:
- Authenticates the client
- Calls
SET NXin Redis for the identity’s lock key with a 30-second TTL - If the lock is already held → rejects with
DuplicateIdentityError - If the lock is acquired → the connection is the lock
The client does not call a separate lock API. There is no “acquire lock then connect” sequence. The connection attempt is the lock attempt.
Heartbeat via Lock Refresh
Section titled “Heartbeat via Lock Refresh”While the stream is open, a background goroutine in the gateway refreshes the Redis lock TTL every 10 seconds. The TTL is 30 seconds, so a gateway crash gives a 30-second grace period before the lock auto-expires and the identity becomes available for reconnection.
Clients do not send heartbeat pings. The TCP keepalive and gRPC stream liveness are the heartbeat. If the client process crashes, the TCP connection closes, the gateway detects the EOF, and lock cleanup happens immediately — not after a TTL expiry.
Lock Release on Disconnect
Section titled “Lock Release on Disconnect”When the stream closes (client disconnect, network failure, gateway shutdown, admin force-disconnect), the gateway:
- Calls
DELon the Redis lock key - Decrements the workspace quota counter
- Updates any associated orchestration task state
- Writes the audit log entry
The lock is released synchronously as part of disconnect cleanup. Other gateway instances can immediately see the identity as available in Redis.
Exclusivity Guarantee
Section titled “Exclusivity Guarantee”Two clients cannot hold the same identity simultaneously. The Redis SET NX operation is atomic. Even across a horizontally scaled cluster with many gateway instances, only one SET NX call wins. The second caller receives DuplicateIdentityError.
Principal types and their uniqueness constraints:
| Principal | Uniqueness |
|---|---|
| Agent | One connection per workspace + implementation + specifier |
| Task (Unique) | One connection per workspace + implementation + unique_specifier |
| Task (Non-Unique) | Multiple connections allowed; server assigns unique IDs |
| User | One connection per user_id + window_id (multiple tabs = multiple window IDs) |
| Workflow Engine | One active connection |
| Metrics Bridge | One active connection |
| Orchestrator | One connection per implementation + specifier |
| Service | One connection per implementation + specifier |
Session Resume
Section titled “Session Resume”When a client reconnects after a brief network partition, the previous lock may still be held in Redis (within the 30-second TTL). The client can include a resume_session_id in its InitConnection message to request atomic lock takeover.
The gateway executes a Lua script that atomically checks whether the existing lock belongs to the resuming session and replaces it if so. This prevents a race between the old lock expiring and a different client acquiring it.
The ConnectionAck response includes a resumed: true flag when this succeeds.
Failure Scenarios
Section titled “Failure Scenarios”Client Crashes
Section titled “Client Crashes”The TCP connection closes immediately. The gateway detects EOF, runs disconnect cleanup, and releases the lock. The identity is available for reconnection within milliseconds.
Network Partition (Brief)
Section titled “Network Partition (Brief)”The TCP connection hangs. The gateway’s keepalive probe eventually times out and closes the stream. The lock TTL (30s) provides a window during which the client can reconnect with resume_session_id and atomically take over the lock without a gap.
Gateway Crashes
Section titled “Gateway Crashes”The gateway process dies without running cleanup. Redis lock TTLs expire after 30 seconds. Clients reconnect to a different gateway instance. RabbitMQ Streams preserve message offsets so no messages are lost.
Thundering Herd on Restart
Section titled “Thundering Herd on Restart”When restarting a container before the previous lock expires, clients receive DuplicateIdentityError. All three SDKs support a RetryOnDuplicate / retry_on_duplicate option that treats this error as recoverable and retries with backoff until the old lock expires (~30 seconds).
Operational Implications
Section titled “Operational Implications”- No separate heartbeat endpoint. There is no
/api/heartbeator keepalive message. If you are looking for one, you are working with a different system. - No zombie sessions. A session in the active sessions list is live by definition. Dead sessions expire within 30 seconds.
- Lock visibility is binary. Either a client holds the identity or it does not. There is no “lock degraded” or “lock suspect” state.
- Admin disconnect is immediate. The
DELETE /api/connections/{session_id}endpoint sends aFORCE_DISCONNECTsignal, which causes the stream to close and the lock to release synchronously.
Comparison to Alternatives
Section titled “Comparison to Alternatives”| Approach | Lock | Heartbeat | Channel |
|---|---|---|---|
| Aether | Implicit (stream = lock) | Implicit (stream liveness) | gRPC stream |
Redis SET NX + heartbeat timer | Explicit SET NX | Separate timer/cron | Separate |
| ZooKeeper ephemeral nodes | Ephemeral node | Session keepalive | Separate |
| etcd leases | Lease grant | Lease keepalive | Separate |
Aether’s approach is simpler operationally because the failure modes for lock, heartbeat, and channel are identical — they all reduce to “is the TCP stream alive.”
See Also
Section titled “See Also”- Identity Model — principal types and their uniqueness rules
- Architecture — how the gateway acquires and refreshes locks