Chat Application Architecture, Part 1 of 4: A Minimal Backend
A practical chat application architecture for the first version: one store, one send path, and the few rules you should not skip before you add scale.

A chat application architecture is easy to overbuild on day one and easy to under-specify where it actually hurts. This series is for solution architects and the engineers who have to draw the boundaries: what is in the first context, what is a later service, and what you refuse to put on the phone.
The first version does not need a broker, a push fleet, or end-to-end encryption. It does need a clear answer to one question: when a user sends a message, where is the durable copy, and how does the other person get it?
Here we stay with a minimal messaging app architecture you can defend in a design review. Part 2 covers the services you add when that design breaks. Part 3 is which of those tasks need a message broker, using RabbitMQ. Part 4 covers libraries and products that replace pieces of the pipe.
What you are actually building
A chat is not a generic CRUD API with a body field. It is an append-only log per conversation, plus a way to learn that the log grew.
Even a primitive backend has three objects:
| Object | Minimum fields | Why it exists |
|---|---|---|
| User | id, auth identity | Who may send and read |
| Conversation | id, member ids | The boundary of access |
| Message | id, conversation id, sender, body, server time, client id | The log entry |
Membership is not a nice-to-have. Every read and every send checks that the caller belongs to that conversation. Skip this in the prototype and you will leak history the first time someone guesses an id.
The smallest path that works: HTTP
You can ship a usable 1:1 chat with two endpoints and a poll:
POST /conversations/{id}/messageswrites one row and returns the stored message.GET /conversations/{id}/messages?after={sequence}returns newer rows.
The client polls every few seconds while the screen is open. That is crude, and it is still a correct message delivery model for an internal tool or a pilot with tens of concurrent users.
sequenceDiagram
participant A as Client A
participant API
participant DB as Message log
participant B as Client B
A->>API: POST message
API->>API: membership check
API->>DB: insert next sequence
API-->>A: stored row
B->>API: GET after cursor
API->>DB: rows after sequence
API-->>B: new messages
Postgres (or any transactional database you already operate) is enough. One table, an index on (conversation_id, sequence), and a transaction that allocates the next sequence for that conversation. Global ordering across all chats is unnecessary. Order inside one conversation is the product requirement.
Three rules that belong in version one
Idempotency. Mobile networks retry. The client sends a client_message_id (a UUID it generated). The server stores it with a unique constraint per sender. A retry returns the original row instead of a second bubble.
Server time, client display. Sort and paginate by the server sequence, not by the phone clock. Show the client timestamp only as a hint.
Cursor, not offset. History uses after=sequence (or before= for older pages). Offset pagination drifts when new rows arrive at the top of the thread.
If those three are in place, you can change the transport later without changing the product semantics.
One step up: a single WebSocket process
Polling wastes battery and feels late. The next increment is still small: one process that holds WebSocket connections and, on each successful insert, pushes the new message to sockets whose user is a member.
sequenceDiagram
participant A as Client A
participant P as Chat process
participant DB as Message log
participant B as Client B
A->>P: WebSocket send
P->>DB: insert
P-->>B: hint, if the socket is open
B->>P: GET history after last sequence
Keep the HTTP history endpoint. The socket is a hint that something new exists. The database remains the source of truth. On reconnect, the client calls history with its last sequence and fills the gap. That pattern survives every later redesign.
Operational minimum for that one process:
- Heartbeats, so dead connections do not sit forever.
- Auth on connect (short-lived token), not a long-lived password in the socket URL query that ends up in logs.
- Backpressure: if a client stops reading, do not buffer unbounded outbound messages in memory. Drop the socket and let history catch up.
This is still one machine. It is the right chat backend architecture until you need a second instance.
What you should leave out on purpose
| Tempting feature | Why it waits |
|---|---|
| Group chats with hundreds of members | Fan-out and membership changes dominate the design |
| Read receipts and "typing…" | Extra events, privacy questions, more write load |
| Push when the app is killed | A separate credential store (FCM/APNs) and a retry policy |
| Media, voice, reactions | Object storage, moderation, and a different size of payload |
| Search across history | Another index and a retention policy |
| End-to-end encryption | Key recovery, search, and abuse tooling all change |
| A message broker | Nothing to fan out yet that the database cannot do |
A prototype that includes all of the above is not a prototype. It is an unread architecture document.
Where the minimal design breaks
You will feel the wall in a predictable order:
- Two API instances. A WebSocket on node A does not see a send handled by node B. In-memory "who is online" is a lie.
- The phone sleeps. A closed app does not poll and does not hold a socket. Delivery now depends on push, which this version does not have.
- A second device. The same user has a laptop and a phone. "Delivered" is per device, not per user, and you have not modelled that.
- A busy group. One insert must become N notifications. Doing that inside the request thread stalls the sender.
None of these require a new product. They require extra services around the same log. That is part 2.
If the conversation is "should we own this log at all", skip ahead to SDKs and hosted chat products.
A note on protocols
XMPP and Matrix already define presence, rooms, and sync. Early WhatsApp is widely described as starting from ejabberd (XMPP) and later replacing that stack as the product outgrew the protocol. The lesson for a custom product is modest: a standard protocol can carry version one, but your membership rules, retention, and integrations will still sit in your database. Do not put AMQP or a heavy broker client on the phone. Keep a thin edge (HTTPS and WebSocket) and hide the internals.
What to document before you write more code
- Who may create a conversation, and who may be added later.
- The meaning of "sent": accepted by the server and stored, nothing more.
- The idempotency key and the unique constraint.
- The reconnect rule: history
afterlast sequence, then resume the socket. - The explicit list of features you are not building yet.
That list is the architecture decision record for version one: scope, invariants, and explicit non-goals. Hand it to the people who will implement the log, not as a feature pitch.
Next
Messaging app architecture: the services you add as it grows walks through presence, push, groups, media, and a broker, in the order the pain usually appears.
The same cut is what a custom software delivery team should be asked to build first: a log and a live channel, not a messenger clone. Client sync and, later, push sit with mobile engineering. A broker, when you actually need one, usually lands next to the rest of the Java backend. If you want that boundary reviewed against a system you already run, get in touch.