Chat Application Architecture, Part 1 of 4: A Minimal Backend

A practical chat application architecture for the first version: one store, one send path, and the few rules you should not skip before you add scale.

Chat Application Architecture, Part 1 of 4: A Minimal Backend

A chat application architecture is easy to overbuild on day one and easy to under-specify where it actually hurts. This series is for solution architects and the engineers who have to draw the boundaries: what is in the first context, what is a later service, and what you refuse to put on the phone.

The first version does not need a broker, a push fleet, or end-to-end encryption. It does need a clear answer to one question: when a user sends a message, where is the durable copy, and how does the other person get it?

Here we stay with a minimal messaging app architecture you can defend in a design review. Part 2 covers the services you add when that design breaks. Part 3 is which of those tasks need a message broker, using RabbitMQ. Part 4 covers libraries and products that replace pieces of the pipe.

What you are actually building

A chat is not a generic CRUD API with a body field. It is an append-only log per conversation, plus a way to learn that the log grew.

Even a primitive backend has three objects:

Object Minimum fields Why it exists
User id, auth identity Who may send and read
Conversation id, member ids The boundary of access
Message id, conversation id, sender, body, server time, client id The log entry

Membership is not a nice-to-have. Every read and every send checks that the caller belongs to that conversation. Skip this in the prototype and you will leak history the first time someone guesses an id.

Minimal chat path: client, API, message log, optional socket

The smallest path that works: HTTP

You can ship a usable 1:1 chat with two endpoints and a poll:

  1. POST /conversations/{id}/messages writes one row and returns the stored message.
  2. GET /conversations/{id}/messages?after={sequence} returns newer rows.

The client polls every few seconds while the screen is open. That is crude, and it is still a correct message delivery model for an internal tool or a pilot with tens of concurrent users.

sequenceDiagram
  participant A as Client A
  participant API
  participant DB as Message log
  participant B as Client B
  A->>API: POST message
  API->>API: membership check
  API->>DB: insert next sequence
  API-->>A: stored row
  B->>API: GET after cursor
  API->>DB: rows after sequence
  API-->>B: new messages

Postgres (or any transactional database you already operate) is enough. One table, an index on (conversation_id, sequence), and a transaction that allocates the next sequence for that conversation. Global ordering across all chats is unnecessary. Order inside one conversation is the product requirement.

Three rules that belong in version one

Idempotency. Mobile networks retry. The client sends a client_message_id (a UUID it generated). The server stores it with a unique constraint per sender. A retry returns the original row instead of a second bubble.

Server time, client display. Sort and paginate by the server sequence, not by the phone clock. Show the client timestamp only as a hint.

Cursor, not offset. History uses after=sequence (or before= for older pages). Offset pagination drifts when new rows arrive at the top of the thread.

If those three are in place, you can change the transport later without changing the product semantics.

One step up: a single WebSocket process

Polling wastes battery and feels late. The next increment is still small: one process that holds WebSocket connections and, on each successful insert, pushes the new message to sockets whose user is a member.

sequenceDiagram
  participant A as Client A
  participant P as Chat process
  participant DB as Message log
  participant B as Client B
  A->>P: WebSocket send
  P->>DB: insert
  P-->>B: hint, if the socket is open
  B->>P: GET history after last sequence

Keep the HTTP history endpoint. The socket is a hint that something new exists. The database remains the source of truth. On reconnect, the client calls history with its last sequence and fills the gap. That pattern survives every later redesign.

Operational minimum for that one process:

  • Heartbeats, so dead connections do not sit forever.
  • Auth on connect (short-lived token), not a long-lived password in the socket URL query that ends up in logs.
  • Backpressure: if a client stops reading, do not buffer unbounded outbound messages in memory. Drop the socket and let history catch up.

This is still one machine. It is the right chat backend architecture until you need a second instance.

What you should leave out on purpose

Tempting feature Why it waits
Group chats with hundreds of members Fan-out and membership changes dominate the design
Read receipts and "typing…" Extra events, privacy questions, more write load
Push when the app is killed A separate credential store (FCM/APNs) and a retry policy
Media, voice, reactions Object storage, moderation, and a different size of payload
Search across history Another index and a retention policy
End-to-end encryption Key recovery, search, and abuse tooling all change
A message broker Nothing to fan out yet that the database cannot do

A prototype that includes all of the above is not a prototype. It is an unread architecture document.

Where the minimal design breaks

You will feel the wall in a predictable order:

  1. Two API instances. A WebSocket on node A does not see a send handled by node B. In-memory "who is online" is a lie.
  2. The phone sleeps. A closed app does not poll and does not hold a socket. Delivery now depends on push, which this version does not have.
  3. A second device. The same user has a laptop and a phone. "Delivered" is per device, not per user, and you have not modelled that.
  4. A busy group. One insert must become N notifications. Doing that inside the request thread stalls the sender.

None of these require a new product. They require extra services around the same log. That is part 2.

If the conversation is "should we own this log at all", skip ahead to SDKs and hosted chat products.

A note on protocols

XMPP and Matrix already define presence, rooms, and sync. Early WhatsApp is widely described as starting from ejabberd (XMPP) and later replacing that stack as the product outgrew the protocol. The lesson for a custom product is modest: a standard protocol can carry version one, but your membership rules, retention, and integrations will still sit in your database. Do not put AMQP or a heavy broker client on the phone. Keep a thin edge (HTTPS and WebSocket) and hide the internals.

What to document before you write more code

  • Who may create a conversation, and who may be added later.
  • The meaning of "sent": accepted by the server and stored, nothing more.
  • The idempotency key and the unique constraint.
  • The reconnect rule: history after last sequence, then resume the socket.
  • The explicit list of features you are not building yet.

That list is the architecture decision record for version one: scope, invariants, and explicit non-goals. Hand it to the people who will implement the log, not as a feature pitch.

Next

Messaging app architecture: the services you add as it grows walks through presence, push, groups, media, and a broker, in the order the pain usually appears.

The same cut is what a custom software delivery team should be asked to build first: a log and a live channel, not a messenger clone. Client sync and, later, push sit with mobile engineering. A broker, when you actually need one, usually lands next to the rest of the Java backend. If you want that boundary reviewed against a system you already run, get in touch.