Chat Application Architecture, Part 4 of 4: Which Parts You Can Take as Shortcuts
A build-versus-buy map for chat backends: CPaaS products, open-source servers, and libraries that only solve transport, and when a custom log is still the right core.

Part 1 showed a minimal chat application architecture: a log and a live channel. Part 2 showed how many services appear once that log meets real phones and real groups. Part 3 is the case where you still own the log and add a broker for fan-out, push, and integrations. You do not have to own all of them.
This article is the options paper a solution architect writes before a vendor workshop. It separates three things people lump together as "a chat library": a product you rent, a server you host, and a library that only moves bytes. The right box depends on whether the conversation is the product or a feature inside another system.
Three layers, three buying decisions
| Layer | You need | Typical options | You still own |
|---|---|---|---|
| Transport | Connections, reconnect, fan-out of packets | Socket.IO, Centrifugo, Ably, Pusher, Phoenix Channels, SignalR | The message log, or you accept theirs |
| Messaging product | Threads, receipts, push, moderation UI | Stream Chat, Sendbird, Twilio Conversations, CometChat, Firebase | Membership rules if you mirror them |
| Full workplace or community server | Accounts, rooms, clients already built | Mattermost, Rocket.Chat, Zulip, Element + Synapse (Matrix) | Almost nothing, unless you fork |
A library is not a messenger. Socket.IO will not give you GDPR deletion, attachment scanning, or a defensible read-receipt policy. A CPaaS will, inside its model, and it will charge and constrain you inside that model.
Libraries: use them when the log is yours
Keep your database as the source of truth (as in part 1) and add a realtime component.
| Option | Fits | Watch-outs |
|---|---|---|
| Socket.IO (Node) | One team, one language, fast prototype | Sticky sessions and a Redis adapter once you run more than one node. Easy to treat the socket as the database. |
| Centrifugo | Polyglot backends, serious connection counts | You still write history, auth tokens, and push. |
| NATS or Redis pub/sub | Internal fan-out between your own stateless nodes | Not a client protocol for mobile apps. Put a gateway in front. |
| Phoenix Channels / SignalR | You are already on Elixir or .NET | Same rule: persist first, then broadcast. |
| Ably, Pusher | You want hosted connections without a full chat product | Payload and retention limits. History may be short. |
| LiveKit, mediasoup | Calls and rooms, not text | Do not force chat history through the media server. |
Rule carried from the first article: do not ship a raw AMQP or MQTT client to a consumer phone unless you have a device constraint that HTTPS and WebSocket cannot meet. Gateways exist so mobile bugs stay in a thin protocol.
Products: use them when threads are a feature
Hosted chat SDKs (Stream Chat, Sendbird, Twilio Conversations, CometChat, and similar) sell the part 2 checklist: channels, typing, push templates, unread counts, moderation hooks, and client UI kits.
They are a good fit when:
- Time to a credible UI matters more than a custom log shape.
- Conversations are attached to an object you already have (a booking, a shipment, a class), and you can store that id as channel metadata.
- You can accept the vendor as a subprocessor, including region and retention.
- You will not need server-side search or workflow that their API cannot express.
They are a poor fit when:
- Message history is a system of record for a regulated process, and residency or audit rules fight the vendor's region.
- The thread model is odd (hierarchical tickets, dual-control approval, messages that are also ledger events).
- Unit economics break at your message volume. Price the busy-group case, not the demo.
- You must delete or export on a schedule their API only approximates.
Firebase (Firestore listeners plus FCM) sits in this bucket for mobile-first teams. It is excellent until security rules and fan-out cost become the product. Plan an exit if the chat is strategic: export shape, id mapping, and a deadline.
Amazon Chime SDK messaging and Azure Communication Services are the same trade for teams already committed to that cloud's identity and region story.
Open protocols and self-hosted servers
| Stack | What you get | What you take on |
|---|---|---|
| Matrix (Synapse, Dendrite) + Element | Federated rooms, E2E options, several clients | Ops, database growth, and upgrade discipline |
| XMPP (ejabberd, Prosody) + client libs | A long-lived standard, extensions for MUC and push | Extension soup. Your product semantics may not map cleanly |
| Rocket.Chat, Mattermost | Slack-like workspace, self-hosted | A workplace product, not an embeddable thread in your marketplace |
| Zulip | Topic-based team chat | Same: a destination app, not a widget |
Self-hosting answers data-residency reviews better than a US-only SaaS, and it costs you the on-call load the SaaS hid. Forking the server to match a strange domain model is usually more expensive than a small custom log plus a transport library.
Matrix is the option to name when a stakeholder says "we might need to federate with other organisations later". Federation is a product promise, not a library flag. If you will never federate, do not put it in the target architecture.
Large open-source projects
The rows above are still components or workplace servers you mostly run as they are. Signal and Jitsi are a different shortcut: whole open-source products, with their own clients, protocols, and release trains. You take them when the goal is to start from a working messenger or a working meeting stack, then change the parts that do not match the product.
| Project | What the shortcut covers | What stays yours |
|---|---|---|
| Signal | An encrypted messenger: protocol, clients, and the message path | Product changes on an older codebase: roles, profiles, and how the messenger joins the rest of the system |
| Jitsi | Self-hosted video meetings | Running it beside the chat. Text history stays in the message log, not in the media server |
That split is the one we used on the blockchain-based social ecosystem: open-source Signal as the messenger foundation, Jitsi for video conferencing, and a blog and wallet built around them. The cost showed up as legacy code inside Signal, not as a missing library.
A simple decision
flowchart TD
A[Is the thread the product, with unusual rules?]
A -->|yes| B[Custom log and your own services]
B --> C[Optional transport library]
A -->|no| D[Is a workplace clone enough?]
D -->|yes| E[Mattermost, Rocket.Chat, Zulip, Element]
D -->|no| F[Chat SDK]
F --> G[Keep your user id and business object id]
Forking Signal or embedding Jitsi is the other shortcut from the section above. It is closer to adopting a product than to adding Socket.IO.
Hybrid that works in practice: the vendor or the open-source server carries generic realtime and push, while your backend remains the authority for "may this user see this shipment thread" and for retention. If you cannot enforce authorisation in the vendor with a token you mint, do not integrate it. Leaking history across tenants is the failure mode that ends the project.
Compare this with the broader custom versus off-the-shelf choice. Chat is just a sharp case of the same question: commodity workflow versus a model you must control.
Evaluation checklist
- Region and subprocessors, in writing, for the EU entity that owns the data.
- Export and erase: one user, one tenant, one conversation. Timed in a trial, not in a slide.
- Idempotency and ordering guarantees. Ask what a retry does.
- Behaviour at your largest channel size, with push enabled.
- UI kit versus headless API. A kit speeds a demo and fights a design system.
- Lock-in: can you read history without their SDK if the contract ends?
- Calls, if you need them: confirm they are a separate product and a separate bill.
What we would not outsource
Even when the pipe is rented, keep these in your system:
- The link between a conversation and your business object.
- The policy for who is a member when that object changes (order cancelled, employee offboarded).
- Retention that matches your privacy notice.
- The audit of staff who open a customer thread.
That boundary is the solution design: a smaller custom software context around a chat SDK, not a second messenger. Mobile engineering still owns token refresh, background limits, and the local queue. If the chosen design is a broker plus your own log, Java backends are a common place to run it.
Get in touch if you want a second architect to pressure-test the box you picked: region, a non-standard thread model, or an existing system of record. The series is the map. The design starts from which box you are actually in.