Get in touchcorey@spiritdevs.com

Principal Engineer, Web and Mobile Platform Architecture · Corporate Interactive · Sydney, Australia

Theme
github.com/coreybainSnapshot 09 OCT 2026 · 00:08 UTC
Direct line

Start a conversation.

Send a note to corey@spiritdevs.com. Add a brief, job specification, or other context if it helps.

AttachmentsUp to 3 files · 4 MB combined

Submitting stores the message and nothing else — no queue in front of it, no autoresponder, no list to be added to.

06 / 06Writing

Real-Time Is Easy Until It Drops

Fri 18 Sep 20264 min readCorey Baines

A message can reach the server while its sender is still waiting. Building Pathway means dealing with that gap.

  • Real-time systems
  • Software engineering
  • Pathway
A message reaches a server while the acknowledgement path back to the laptop is interrupted.

On this page

  1. The connection drops at an awkward moment
  2. A connection is only part of the state
  3. Reconcile before clearing the draft
  4. A slower response can arrive last
  5. Give uncertainty a visible state
  6. Test the interruption

Share

The connection drops at an awkward moment#

Imagine sending a message, then losing your connection before the interface confirms it.

There are several possible explanations. The message never reached the server. It reached the server and is waiting to run. Or the server accepted it, but the response never made it back.

From the client, those situations can look identical.

Send it again and you might duplicate the work. Clear it and you might lose the only copy. Leave the spinner running and the person has to guess whether anything is happening.

These are the kinds of cases I've been working through in Pathway, where messages can launch work on another machine.

A hypothetical sequence: the server accepts a message, its acknowledgement is lost, and the client checks live state after reconnecting.

A connection is only part of the state#

A connected socket tells you that the client and server can communicate. It doesn't tell you whether a particular message was accepted or whether the client's view has caught up.

The interface needs to distinguish a local draft, a pending send, queued work and a message confirmed by server state.

Those distinctions affect the controls. A local draft can be edited directly. Queued work needs queue actions. Work that has already started needs controls that act on the running task.

Treating all of that as “sending” makes recovery hard to explain and easy to get wrong.

Reconcile before clearing the draft#

Pathway preserves information about a pending draft send so it can be checked after a reload.

That check waits for live state. A cached conversation list isn't enough evidence to decide that a send failed.

The reconciliation code checks accepted threads and queue state before deciding what to do with the local draft. It also protects the handoff from the draft view to the actual conversation, so the draft doesn't disappear before the message becomes visible.

If a restored pending send isn't represented in the accepted or queued state once those checks are ready, the code can restore its content to the composer.

That recovery also has to preserve anything the person typed while the original send was underway. Clearing the old message must not clear their new one with it.

A slower response can arrive last#

Reconnect handling also means dealing with information arriving out of order.

One Pathway change moved server replicas from frequent polling to a live subscription for the latest sync version. A slower HTTP check remains as a recovery path.

That creates a race: an HTTP request can start first, return later, and carry an older version than the subscription has already delivered.

The transport rejects duplicate or regressing versions within the same authorisation epoch. Changes to that epoch are handled separately because access changes matter too.

Without that ordering rule, adding a recovery path could move the replica backwards.

Give uncertainty a visible state#

“Disconnected” and “failed” mean different things to someone waiting for their work.

A disconnected client may be unable to confirm the result. A failure means it has evidence that something went wrong. The available actions should reflect that difference.

The UI should explain when it's checking a send, when work is queued and when the person can safely edit a recovered draft. It also needs to stop showing recovery once the work has settled.

These details are part of the behaviour people rely on when they switch devices, close a laptop or return to an old tab.

Test the interruption#

A useful recovery test puts the interruption at a specific boundary: after sending but before confirmation, during queue loading, or while a newer update overtakes an older response.

Pathway's draft tests include cached state, a queue that hasn't finished loading and follow-up text that must survive reconciliation.

For this part of the system, the result I care about is concrete: after reconnecting, the person can see what happened to their message, and the text they still need is there.

02Keep readingAll writing
An annotated screenshot, a recording and diagnostic logs collected into one support report.
Newer post

Bug Reports Should Bring Context

Fri 25 Sep 2026