Chat Application Architecture, Part 2 of 4: Services You Add as It Grows
Which services a messaging app architecture needs after the first backend: presence, push, groups, media, and a broker, mapped to the failure that forces each one.

The minimal chat application architecture in part 1 is a conversation log plus a way to hear about new rows. Solution architects hit the same wall in design reviews: a second instance, a sleeping phone, a group that is no longer a loop in the request thread. Each failure is a reason to add a service, not a reason to redraw the whole system.
This messaging app architecture guide lists those services in the order teams usually need them. Part 3 is the broker tasks, worked on RabbitMQ. Part 4 is the alternative: rent or embed someone else's pipe instead of growing yours.
Add services when a symptom shows up
| Symptom | Service you add | What stays the same |
|---|---|---|
| Second app server, missed live messages | Connection registry + pub/sub between nodes | Message log in the database |
| App killed, user sees nothing until reopen | Push gateway (FCM, APNs) | History catch-up on open |
| "Seen" and typing indicators | Ephemeral event channel | Do not store typing in the message log |
| Group of dozens or hundreds | Fan-out worker, membership version | Per-conversation sequence |
| Laptop and phone disagree | Device registry + per-device cursor | One logical message id |
| Photos and files | Media pipeline (object storage, virus scan, CDN) | Message row points at an attachment id |
| Restarts drop in-flight work | Transactional outbox + broker | At-least-once plus idempotency |
| Abuse, legal hold, GDPR delete | Retention jobs, audit of admin actions, report/block | Authorisation on every read |
If a feature is not on this list yet, you probably do not need Kafka.
1. Connection registry and presence
Once more than one process accepts WebSockets, "who is connected" cannot live in process memory. A small registry (often Redis, or a table plus pub/sub) maps user_id → connection ids → node.
On send:
sequenceDiagram
participant API
participant DB as Message log
participant Bus as Internal bus
participant Node as Socket node
participant Push as Push worker
API->>DB: commit message
API->>Bus: conversation id and sequence
Bus->>Node: nodes that hold a member socket
Bus->>Push: devices with no open socket
Node-->>Node: client loads the row, or takes the payload
- Commit the message.
- Publish a lightweight event (
conversation_id,sequence) on an internal bus. - Each node that holds a member's socket pushes that event.
- The client fetches the row, or the event already carries the payload if you accept the size.
Presence ("online") is a projection of that registry with a timeout. Treat it as a hint. Do not use presence as the only delivery mechanism. Offline users never appear in the registry, and they still must receive the message later.
2. Push, because sockets die
iOS and Android will suspend your process. Offline message sync then has two parts:
- A durable inbox (the same log).
- A push that says "open the app" or, within payload limits, shows a preview.
Collapse notifications by conversation so a busy thread does not become 40 banners. Store device tokens per user, delete them when the vendor says the token is dead, and never block the send request on the push provider. Push is best-effort. The log is not.
3. Receipts without pretending they are messages
Delivery states that product cares about:
| State | Who sets it | Stored? |
|---|---|---|
| Accepted | Server, after commit | Yes, it is the row |
| Pushed | Push worker | Optional metric, not a user-facing guarantee |
| Delivered to device | Client ack | Per device, if you promise the tick |
| Read | Client, when the thread is visible | Per user, with a privacy review |
"Read" is personal data. In a workplace tool you may show it. In a consumer or health-adjacent chat you may only store a watermark (last_read_sequence) and avoid broadcasting a precise timestamp to the sender. Decide that before the UI draws two blue ticks.
Typing and presence events are ephemeral. Put them on the socket bus and set a short TTL. If they land in the message table, you will paginate through noise.
4. Groups and fan-out
A group message is still one row in the conversation log. The expensive part is notifying members.
Do not loop "send push + socket" inside the HTTP request once membership is large. Commit the row, then a worker reads membership and fans out. Membership changes need a version: a user removed at sequence 500 must not receive sequence 501, and a user who joins late should not automatically receive the entire history unless the product says so.
Ordering stays per conversation. Partition or shard hot conversations by conversation_id if one room dominates CPU. You still do not promise one global order across rooms.
Idempotency from part 1 still applies. Fan-out retries will otherwise duplicate notifications even when the log row is unique.
5. Multi-device sync
One user, many endpoints. Model it explicitly:
devices(token, platform, last seen).- A cursor per device, or one cursor per user if you only promise "read" at user level.
On reconnect each device asks for sequence > cursor. The server does not try to remember the WebSocket backlog from yesterday. That is how offline message sync stays boring and correct.
Conflict policy for edits and deletes can stay server-authoritative (last write wins on edited_at) until you have a real offline editor. CRDTs are justified for collaborative text, rarely for a chat bubble.
6. Media as a side pipeline
The message body stays small. Upload flow:
- Client requests an upload slot (authenticated, size-capped, type-allowlisted).
- Bytes go to object storage, preferably not through your API process.
- A scanner runs.
- The client sends a message that references
attachment_id. - Readers get a short-lived URL. The CDN caches bytes, not access checks. Access checks stay on the API that mints the URL.
Voice and video calls are a different system (media servers, NAT traversal). Do not hide them inside the chat log. Link a call id from a message if the product needs a history entry.
7. Broker, outbox, and delivery guarantees
A broker (RabbitMQ, NATS, Kafka) earns a place when several workers must react to "message committed": fan-out, push, search indexing, webhooks into a CRM, analytics. It does not replace the database.
Practical guarantee: at-least-once delivery plus idempotent consumers. End-to-end exactly-once is expensive and rarely what the UI needs. The user needs one bubble, which you already get from client_message_id.
Use an outbox so the database commit and the event publish cannot diverge: the same transaction writes the message and an outbox row; a publisher relays the outbox to the broker. Failed consumers go to a dead-letter queue with a replay runbook.
Kafka fits a high volume of ordered events per conversation key. RabbitMQ or NATS fits task-shaped work (push, webhooks) with simpler operations. Pick from the failure mode, not from a logo slide. Teams that already run enterprise Java services often already have one of these and should not add a second bus for chat alone.
Which of these tasks actually need a broker, and how RabbitMQ takes each one, is part 3.
8. Security and abuse as features, not a layer you paint on
By the time you have groups and media, the security bar is no longer "we use TLS".
- TLS everywhere. Pin certificates only if you can rotate them without bricking the app.
- Short-lived access tokens on the socket.
- Authorisation on every history page, not only on connect.
- Encryption at rest for the database and the bucket, with backups included.
- Retention and deletion that match the privacy notice (export, erase, legal hold).
- Rate limits per user and per conversation. Report and block as first-class membership states.
- Admin audit logs that record who exported or deleted, not a second copy of every private body.
End-to-end encryption is a product decision. It breaks server-side search, many moderation tools, and simple password reset. Offer it when the threat model requires it, and budget for key management. Do not badge a TLS-only chat as end-to-end.
For regulated payloads, the same habits as in a fintech security checklist apply: data map, subprocessors, and a deletion path you have actually tested.
9. What to measure
- Send path latency (accept to commit).
- Live fan-out delay (commit to socket push) at p95.
- Push failure rate and invalid-token rate.
- Depth of the outbox and the dead-letter queue.
- Reconnect storms after a deploy (connection churn).
- Hottest conversation ids.
Load-test the fan-out, not only the empty WebSocket handshake. A marketing blast into a large group is the incident you can rehearse.
When to stop adding services
Stop when the next requirement is "be Slack" and your domain is a narrow thread inside another product (order chat, care team, marketplace). Hosted SDKs and open-source servers exist for the generic pipe. Custom logic should remain the membership rules, the retention policy, and the integration with your system of record. That split is part 4.
If the domain model is yours, write the roadmap as phases with a trigger for each phase: log and socket, then push, then groups, then media. Do not schedule the broker because a reference diagram had one. Custom software delivery follows that sequence. Mobile clients own the cursor and the push token lifecycle. Get in touch if you want the phase triggers checked against a traffic sketch and an existing integration landscape, not against a feature list.