Chat Application Architecture, Part 3 of 4: Which Tasks Need a Message Broker (RabbitMQ)

Which chat tasks need a message broker, and how RabbitMQ takes each one: realtime fan-out, push, and integrations, without becoming the message history.

Chat Application Architecture, Part 3 of 4: Which Tasks Need a Message Broker (RabbitMQ)

Part 2 adds services around the log. Not every service needs a message broker. This part lists the tasks that do, then shows how RabbitMQ takes each task. It is not a reason to add a broker in version one. Part 4 is the alternative: rent the pipe instead of running this bus.

Which tasks need a broker

A broker is justified when more than one kind of work must happen after a message is stored, and that work must not sit inside the request that wrote the row.

Task Why the request thread cannot own it
Realtime fan-out One committed message must wake every socket node that holds a member. Doing that in the HTTP handler ties send latency to connection count.
Push FCM and APNs are slow and fail independently. The sender should not wait, and a push retry must not replay the live socket event.
Integrations Webhooks, search indexing, and analytics fail often and must not delay the bubble in the UI.
Handoff across a crash The row can commit and the process can die before workers hear about it. The handoff has to be recoverable.

Presence, typing, and a single-process WebSocket hint are not on this list. They are ephemeral or local. Part 1 covers that case: no broker yet.

The broker is not the chat

Two stores get confused in design reviews:

Store Holds Survives Ordered by
Message log (database) The bubbles users can scroll Yes Sequence per conversation
Broker A unit of work for a worker Until acked, or per queue policy Queue order, which is not chat order

If RabbitMQ is the only copy of a message, a queue purge or a short TTL deletes history. Keep the log in the database. Publish an event that names the row (message_id, conversation_id, sequence). Workers that need the body load it, or the event carries a size-capped preview for push.

RabbitMQ topology beside the chat log

How RabbitMQ takes those tasks

One exchange, one queue per task. Do not create a queue per user, per device, or per conversation. That looks like isolation and becomes an operations problem at tens of thousands of queues: memory, alarms, and reconnect storms that redeclare them.

RabbitMQ mapping:

Queue Consumer Why it is separate
realtime Socket nodes, or a dispatcher in front of them Fast, in cluster. Must not wait on Apple or Google.
push Push worker Slow external HTTP. Own prefetch and own retry.
integrations Webhooks, search index, analytics Failures here must not delay the tick in the UI.

Bind all three to the same routing key, for example message.committed, on a topic exchange chat.events. Competing consumers on one queue share that task. A second copy of the event is a second queue, not a second publish from the API. Realtime, push, and integrations then fail and scale on their own.

flowchart LR
  API[API commit] --> OB[Outbox]
  OB --> EX[Exchange chat.events]
  EX --> RT[Queue realtime]
  EX --> PS[Queue push]
  EX --> IN[Queue integrations]
  RT --> DLX[Dead-letter exchange]
  PS --> DLX
  IN --> DLX

Handoff: outbox, not a publish inside the request

The same database transaction writes the message and an outbox row. A publisher process reads the outbox, publishes with publisher confirms, then marks the row sent. A crash between commit and confirm retries the publish. Consumers must tolerate the duplicate. That is the recoverable handoff from the task table.

Acks, prefetch, and the dead letter

Ack when the side effect is done, not when the handler starts. A push worker that acks before FCM answers will drop the notification on a process kill.

Set prefetch for the slowness of the side effect. A push worker waiting on a vendor should not hold hundreds of unacked messages. A realtime dispatcher that only looks up sockets can run a higher prefetch.

Failed messages go to a dead-letter exchange, not into an infinite requeue. Retry a few times with backoff for timeouts. Poison messages (bad payload, revoked token you will never fix by retrying) land in a dead-letter queue with a replay runbook. Alert on depth, as in the metrics list in part 2.

Quorum queues are the usual choice when the broker itself must survive a node loss. Classic mirrored queues are the legacy option. Either way, durability of the queue does not make it the system of record.

Ordering is still the database sequence

RabbitMQ can keep order inside one queue only while a single consumer runs and nothing is requeued. The moment you scale the push queue to three workers, two notifications for the same conversation can leave out of order.

That is acceptable if clients apply sequence from the log and ignore a late hint. It is not acceptable if the worker writes the bubble. Workers notify. They do not invent history.

Do not partition RabbitMQ by conversation_id to "fix" this. You will be back to a queue per hot conversation. If you need the bus itself to be the ordered log, that is a different design: Kafka (or similar) with the conversation id as the partition key, and consumers that materialise read models. Pick one source of truth.

What not to put on the phone

AMQP and MQTT clients on iOS and Android were a common shortcut in older messenger write-ups. The failure mode is a heavy client protocol, background limits, and a broker exposed to the internet. Keep a thin edge: HTTPS and WebSocket, tokens, membership checks. The broker stays on the private network. Part 4 makes the same cut for hosted realtime products.

Presence and typing do not belong in a durable queue. They are ephemeral and already covered by the connection registry. A RabbitMQ message for "user is typing" will outlive the typing.

When RabbitMQ is the wrong bus

Need Prefer
Task fan-out you already operate on RabbitMQ Stay. Do not add Kafka beside it for chat only.
Ephemeral hints between socket nodes, no replay NATS or Redis pub/sub is enough. The log is still the database.
Long retention, replay, partition order per conversation as the log Kafka-style log. Do not also treat Postgres as a second history.
One team, one Node process, no second worker yet No broker. Part 1.

Decisions to write down

  • The log is the database. The broker names rows.
  • Exchange, routing key, and the queue list. Explicitly: no queue per user.
  • Outbox plus publisher confirms.
  • Ack point, prefetch, and dead-letter policy per queue.
  • Clients reconcile by sequence. Workers do not reorder the thread.
  • The phone never speaks AMQP.

That is the broker slice of a custom chat backend. The Java services that already run RabbitMQ for other domains are in Java and enterprise backends. Push tokens and background limits stay with mobile engineering. If the open question is whether this bus should exist at all, get in touch with the worker list, not with a vendor logo.