Why your 3CX trunk keeps dropping, and how to find out before your customers do
A SIP trunk that unregisters does not throw an error anyone sees. The phones simply stop ringing, and you find out when somebody walks over to your desk.
Almost every "the phones are down" call turns out to be one of a handful of things. The frustrating part is not diagnosing it — it is that nobody knew until a customer could not get through.
Here is what actually causes a 3CX trunk to unregister, how to tell which one you are looking at, and why the fix is usually monitoring rather than configuration.
The five usual causes
1. The provider dropped the registration. By far the most common. Your SIP provider restarts something, or ages out a registration, and 3CX does not re-register cleanly. Nothing changed on your side, which is exactly why it is so confusing to debug after the fact.
2. The registration interval is fighting a NAT timeout. If your firewall's UDP timeout is shorter than the trunk's re-registration interval, the pinhole closes before 3CX refreshes it. Calls out still work. Calls in vanish. This one is nasty because it looks intermittent.
3. The public IP changed. Dynamic WAN address, or a failover to a secondary link, and the provider is still sending to the old address.
4. Certificate or TLS problems. On a TLS trunk, an expired intermediate certificate will drop registration with no obvious warning. Worth checking alongside your SSL certificate expiry, which fails the same silent way.
5. The 3CX service itself stopped. Less common, but if the SIP service is not running, nothing registers. This is why service state is worth watching separately from trunk state.
Telling them apart
The single most useful thing is a timeline. One trunk down while others stay up points at the provider. Every trunk down at once points at your network, your IP or your 3CX. A trunk that drops at a suspiciously regular interval is almost always the NAT timeout.
That is difficult to reconstruct from memory and easy to read off a history:
- Did it drop alone, or with everything else?
- Does it recur on a schedule?
- Did extensions drop at the same moment, or only trunks?
- Was there a version or licence change around the same time?
Without a record, you are guessing. With 30 days of trunk history, the pattern is usually obvious in about ten seconds.
The real problem is the gap
Take a trunk that fails at 4:50pm on a Friday. If nobody calls in over the weekend, you find out at 8:30am Monday — and by then you have lost two and a half days of inbound calls with no idea they were missed.
The technical fault might be five minutes to fix. The damage is entirely in the gap between it breaking and you knowing.
Most of the outages we see are short. The expensive ones are the short outages nobody noticed for three days.
What to actually monitor
Trunk state alone is not enough, because a trunk can be registered while the system behind it is unusable. The set worth watching together:
- Trunks up, as a proportion. With two trunks, alert below 50% so one failing provider trunk is caught.
- Extensions registered, which catches network problems the trunk state misses.
- 3CX services running, which catches the case where the box is up but the software is not.
- Licence expiry, because an expired licence takes calls down just as effectively as a dead trunk. See 3CX licence expiry for what that actually looks like.
- CPU, memory and disk, because a full disk takes everything with it.
CommsAlert's 3CX monitoring tracks all of these on a schedule you set, with a separate alarm threshold for each. Nothing is installed on the 3CX server — it connects through a dedicated API key.
Setting sensible thresholds
The most common mistake is setting thresholds so tight that you train yourself to ignore the alerts.
| Alarm | A sensible starting point | Why |
|---|---|---|
| Trunk alert | 50% with two trunks | Catches one provider failing without firing on both |
| Extension alert | Well below your daytime figure | Extensions legitimately drop after hours |
| Service alert | 90% | Catches a stopped service without noise |
| Disk usage | 90% | Enough warning to act before recordings fill the disk |
| Licence expiry | 60 days warning, 30 critical | Time to raise a PO before it bites |
Extensions are the one people get wrong. If you have 40 extensions and set the alarm at 90%, you will get an alert every evening as people shut down for the day. Set it against your realistic overnight floor, not your 10am peak.
Telling everyone else
Once you are being alerted, the next problem is the phone call from every user asking whether it is "just them". A public status page answers that without involving you — it shows the current state and 90 days of history on its own link, which you can share with staff or customers.
In short
Trunks drop. That is normal, and largely outside your control. What is inside your control is how quickly you know — and whether you can see enough history to work out which of the five causes you are actually dealing with.
If you want to see this on your own systems, the free three month trial covers 10 phone systems with no credit card and no obligation.
Know the second a trunk drops
CommsAlert checks every trunk on every 3CX system you run, and alerts you by email and SMS the moment the proportion up falls below the threshold you set.