# Trunk Registration Recovery & Boot Ordering — Fix & Verification

**Version:** v1.1.197
**Applies to:** all V-Connect PBX deployments (Asterisk 22)
**Origin:** Edge Financial incident, 2026-07-20 (no inbound/outbound calls all day after power outages)

---

## 1. What went wrong (root cause)

Two hard power-offs at the client site caused two reboots. After the second reboot:

1. Asterisk started **before MariaDB** was ready → ODBC `SQLConnect` errors; realtime PJSIP config (trunks) may not load cleanly on a cold boot.
2. The site's internet was still coming back up, so the SIP trunk registration to the carrier got **"No response" 10 times**, hit `max_retries`, and Asterisk **permanently stopped registering**.
3. With no active registration, the carrier had nowhere to route calls → **inbound AND outbound dead** for the rest of the day, until a reboot/manual re-register.

The key platform gap: **pjsip does not self-recover once `max_retries` is reached.** Nothing re-armed the registration when connectivity returned.

---

## 2. What was fixed (v1.1.197)

Both fixes ship via `pbx-update.sh` — **no manual client action required** to install them.

### Fix 1 — Trunk registration watchdog (the core fix)
- New command: `pbx:trunk-register-watchdog`
- Scheduled **every minute** (`routes/console.php`, `withoutOverlapping`).
- It reads `pjsip show registrations` and, for any trunk **not** `Registered`, re-arms it with `pjsip send unregister/register <id>` — which **resets the retry counter** and starts a fresh attempt.
- Result: after a power cut / network blip / `max_retries` stop, the trunk comes back **automatically within ~60 seconds** of connectivity returning. No reboot, no manual step.
- **Safe by design:** healthy and in-flight states (`Registered`, `Request Sent`, `Auth Sent`) are **left untouched**, so it never flaps a working trunk. It also catches registrations that are missing entirely from the live list.

### Fix 2 — Asterisk waits for MariaDB on boot
- `pbx:deploy` installs a systemd drop-in: `/etc/systemd/system/asterisk.service.d/10-wait-mariadb.conf`
- Orders Asterisk `After=` MariaDB + `network-online.target`, with a **bounded** `ExecStartPre` readiness wait (~60s max).
- On a cold boot Asterisk no longer starts before the DB is ready, so realtime trunk config loads correctly.
- Uses `After=`/`Wants=` (**not** `Requires=`) so a MariaDB restart can never take Asterisk down; the wait always exits after ~60s so Asterisk still starts even if the DB is genuinely down (ODBC reconnects afterwards).

---

## 3. How to deploy on a client box

```bash
sudo /var/www/html/scripts/pbx-update.sh
```

This pulls v1.1.197, installs the watchdog schedule, and writes the Asterisk boot-ordering drop-in. No Asterisk restart is required for the watchdog to start working (it runs from cron/scheduler). The boot-ordering drop-in only affects the **next** boot.

---

## 4. Verification checklist (for the PBX team)

Run these on an updated box to confirm both fixes are active.

### 4.1 Confirm the version
```bash
cd /var/www/html && git describe --tags
# Expect: v1.1.197 (or later)
```

### 4.2 Watchdog command exists
```bash
php /var/www/html/artisan list | grep trunk-register-watchdog
# Expect: pbx:trunk-register-watchdog  ...
```

### 4.3 Watchdog is scheduled (every minute)
```bash
php /var/www/html/artisan schedule:list | grep -i trunk-register-watchdog
# Expect a line showing the command running every minute (* * * * *)
```

### 4.4 Watchdog is SAFE on healthy trunks (no-op test)
With all trunks currently `Registered`, run it manually — it should do **nothing** (no output, no re-registration):
```bash
sudo php /var/www/html/artisan pbx:trunk-register-watchdog
# Expect: NO "re-registered" lines. Silence = healthy trunks left alone.

# Confirm trunks are still registered and untouched:
asterisk -rx "pjsip show registrations"
# Expect: Status = Registered for each trunk
```

### 4.5 Watchdog RECOVERS a down trunk (live recovery test)
> Do this during a quiet window — it briefly unregisters/re-registers one trunk.

```bash
# Pick a trunk id from the registrations list, e.g. ECN123 / reg-ecn:
asterisk -rx "pjsip send unregister reg-ecn"     # simulate a dropped registration
asterisk -rx "pjsip show registrations"          # should now show it NOT Registered

# Wait up to ~60s for the scheduled watchdog (or run it manually):
sudo php /var/www/html/artisan pbx:trunk-register-watchdog
# Expect: "• re-registered trunk 'reg-ecn' (was ...)"

asterisk -rx "pjsip show registrations"           # should be Registered again
```

### 4.6 Boot-ordering drop-in installed
```bash
ls -l /etc/systemd/system/asterisk.service.d/10-wait-mariadb.conf
systemctl show asterisk.service -p After | tr ' ' '\n' | grep -iE "mariadb|mysql"
# Expect: mariadb.service (and/or mysql/mysqld) listed in After=

systemctl show asterisk.service -p ExecStartPre | grep -o "mysqladmin ping"
# Expect: mysqladmin ping  (the readiness wait is present)
```

### 4.7 Watchdog activity is logged when it acts
```bash
grep -i "Trunk registration watchdog" /var/www/html/storage/logs/laravel*.log | tail -5
# Shows each time a trunk was re-armed (only appears when a trunk was down)
```

---

## 5. Simulating the full incident (optional, thorough test)

To prove end-to-end recovery like the real incident:

1. Confirm trunk is `Registered`.
2. Block the carrier IP briefly to force registration to fail past max_retries:
   ```bash
   # 154.119.162.123 = the ECN carrier in the incident; use the box's real carrier IP
   iptables -I OUTPUT -d 154.119.162.123 -j DROP
   ```
3. Force a re-register and watch it fail repeatedly, then stop:
   ```bash
   asterisk -rx "pjsip send unregister reg-ecn"
   asterisk -rx "pjsip send register reg-ecn"
   # over the next minutes: asterisk -rx "pjsip show registrations" → Rejected / not Registered
   ```
4. Restore connectivity (simulating the internet coming back):
   ```bash
   iptables -D OUTPUT -d 154.119.162.123 -j DROP
   ```
5. Within ~60s the scheduled watchdog re-arms the registration:
   ```bash
   asterisk -rx "pjsip show registrations"   # → Registered
   ```
   Without the fix, step 5 never happens on its own — the trunk stays dead until a reboot.

---

## 6. Still recommended (client-side, not software)

- **Fit / verify a UPS** with graceful OS shutdown. The v1.1.197 fixes make the platform **recover automatically** after an ungraceful reboot, but a UPS avoids the hard power-offs (and the filesystem/DB corruption risk) in the first place.
- The two hard power-offs on 2026-07-20 had **no graceful shutdown logged** — that points to raw power loss at the site.

---

## 7. Related commits

| Commit tag | Change |
|---|---|
| **v1.1.197** | `pbx:trunk-register-watchdog` (every-minute re-arm) + Asterisk `After=mariadb` boot-ordering drop-in |
