# AMI / Stasis Subscription Leak — Final Fix & Developer Handoff

**Status:** ✅ TRUE ROOT CAUSE found & fixed — legacy `users.conf` (see below)
**Applies to:** all V-Connect PBX deployments (Asterisk 22)

---

## ⚠️ READ THIS FIRST — the actual root cause was `users.conf`

After the AMI-client work below, the leak was STILL occurring on a fully-updated
box (dev: 40 orphaned subs in ~20h, ~2/hr; live boxes worse). Live diagnosis on
Asterisk 22.4.1 found the real cause, and it has **nothing to do with our PHP /
AMI code**:

> The mere PRESENCE of `/etc/asterisk/users.conf` makes Asterisk **leak one
> `stasis/p:manager:core` subscription on every `dialplan reload` /
> `module reload pbx_config`** — even with ZERO AMI clients connected.

Proof (dev box, Asterisk 22.4.1):
- 3 dialplan reloads with `users.conf` present → **+3** subs.
- Reloads with **0** AMI sessions connected → still **+1 per reload** (so it is
  not the dialler, not mod_php, not any PHP AMI client).
- Move `users.conf` out of `/etc/asterisk/` → **8 reloads = delta 0**. Leak gone.
- Emptying the file is NOT enough; it must not exist so the module stops parsing it.

Why it silently killed these boxes: they reload the dialplan constantly — hourly
sync jobs (`pbx:sync-blf-hints`, `pbx:sync-extension-dialplan`,
`pbx:sync-did-normalizer`, `pbx:audit-recording`), **every** extension
add/edit/delete, IVR / time-condition / ring-group changes, and every deploy.
Each reload orphaned one subscription until `stasis/pool-control` hit its 500 cap
and started dropping events → inbound calls failed and CDR stopped. **This is why
even meyers (dialler disabled) kept degrading.**

Confirmed as a known upstream issue: Asterisk community thread
*"Task processor stasis/p:manager:core orphaned after console reload in Asterisk
22.5"* — resolution: `users.conf` was the culprit.

### The fix
`users.conf` is a legacy chan_sip/chan_iax2 sample. These deployments are
PJSIP + realtime (`ps_endpoints`) and nothing reads it (verified: no code or
scripts reference it). `pbx:deploy` now moves it to `users.conf.disabled` (step
*"Legacy users.conf (stasis leak)"*), which stops all new leaks.

Manual one-liner for an urgent box:
```bash
sudo mv /etc/asterisk/users.conf /etc/asterisk/users.conf.disabled
asterisk -rx "dialplan reload"
# clear subscriptions already orphaned (a reload can't reap them):
asterisk -rx "core show channels count"   # proceed only if 0 active
asterisk -rx "core restart when convenient"
```

> The AMI-client `Events: off` work described below is still correct and stays in
> place (read-only clients should not subscribe to the event firehose, and it
> prevents a *separate* leak on unclean teardown) — but it was **secondary**. The
> `users.conf` removal is the fix that actually stops the recurring flood.

---

## Secondary hardening (still valid): AMI client `Events: off`

> Everything below is the earlier, secondary work on the PHP AMI clients. It is
> correct and stays in place, but it was NOT the primary cause — see the
> `users.conf` section above for the fix that actually stops the flood.

### Background (the secondary AMI-client leak)

A PHP AMI session that logs in **without** `Events: off` also subscribes to the
event firehose and can orphan a `stasis/p:manager:core` subscription on unclean
teardown. This is a real but secondary contributor.

### Root behaviour

When an AMI session logs in **without** `Events: off`, it subscribes to
Asterisk's full event firehose. If that socket is ever torn down uncleanly
(killed process, or an HTTP request ending without a clean disconnect under
mod_php), the subscription is orphaned on Asterisk's side and never reaped.
Under churn these accumulate until the event bus overflows.

---

### The three AMI code paths (all hardened)

The codebase has **exactly three** places that log into AMI. Earlier fixes
covered only two. All three are now correct.

| AMI client | Used by | Fix | Commit |
|---|---|---|---|
| `App\Services\AmiClient` | click-to-dial, dashboard metrics, endpoint status, wallboard | `Events: off` + clean `logoffAndClose()` (reads Goodbye before close) | earlier |
| `App\Services\Asterisk\AmiClient` | `PjsipService` / operator console | `Events: off` | v1.1.182 (`cd48268`) |
| `App\Services\Dialer\AsteriskAmiService` | dialler worker, agent console, pacing | **`Events: off` by default; events made opt-in** | **v1.1.185 (`c5c641b`)** |

The third path was the one still leaking. It is the fix that mattered.

---

## The v1.1.185 fix in detail (the definitive one)

`AsteriskAmiService` was logging in with no `Events:` header, which defaults to
"all events on" — so **every** consumer of it subscribed to the firehose, even
short-lived ones that only run a single action.

The change makes event subscription **opt-in**:

- `AsteriskAmiService::doConnect()` now sends `Events: off` by default.
- Added `subscribeToEvents(bool)` to the `AsteriskAmiInterface` contract (and a
  no-op in `FakeAmiService` for tests).
- Only the long-lived **dialler worker** calls `subscribeToEvents(true)` before
  connecting — it genuinely needs the async events to correlate
  `DialEnd` / `AgentConnect` / `Hangup` via `drainEvents()`.
- Every other consumer (the agent-console originate in `AgentController`, the
  pacing / queue-availability check) stays off. Their action responses
  (QueueStatus rows, Originate result) are delivered regardless of the event
  mask, so nothing breaks.

**Why this is robust:** with `Events: off`, **no subscription is created**, so
even an unclean socket teardown cannot leak something that does not exist. It
does not rely on perfect disconnect handling.

### Login packet (before → after)

```php
// BEFORE (leaked — subscribes to the full event firehose)
$this->write("Action: Login\r\nUsername: {$this->user}\r\nSecret: {$this->secret}\r\n\r\n");

// AFTER (Events: off unless this connection explicitly opted in)
$events = $this->subscribeToEvents ? 'on' : 'off';
$this->write("Action: Login\r\nUsername: {$this->user}\r\nSecret: {$this->secret}\r\nEvents: {$events}\r\n\r\n");
```

---

## Clean disconnect (defense-in-depth, already in place)

Every short-lived consumer already disconnects cleanly — no change needed:

- `AgentController::dial()` wraps the originate in
  `try { ... } finally { $ami->disconnect(); }`.
- `QueueAvailabilityService` uses a context-aware lifecycle: it only
  disconnects if it was the one that opened the connection, so it never closes
  the worker's shared connection.

---

## Deploy hygiene fix (v1.1.186)

A subtle but important gap: **a long-running worker keeps the OLD PHP code in
memory for the life of the process.** A fix deployed by `pbx-update.sh` does not
take effect in an already-running worker until that worker restarts. This is the
mechanism behind "we deployed the fix but it is still leaking."

`pbx-update.sh` already stops/refreshes the dialler worker, but never touched
the Laravel queue worker. Added a step (after the cache rebuild) that runs
`php artisan queue:restart` and hard-restarts the `laravel-queue-worker` unit if
present.

> Note: no queue job in our code uses AMI, so this is **not** itself an AMI-leak
> vector — it is general deploy hygiene so any future code fix reliably reaches
> long-running workers.

---

## Verification on the dev box (vcnew)

- `App\Services\AmiClient`: 5 connect/logoff cycles → **0** new subscriptions.
- `App\Services\Asterisk\AmiClient`: 8 cycles → **0** new subscriptions.
- `AsteriskAmiService` in default (Events: off) mode: 12 connect/action/disconnect
  cycles → **0** new subscriptions.
- Restarted Asterisk (0 active calls): subscription count dropped from a climbing
  73 → **0**, and held flat at 0 over several minutes **with the dialler worker
  connected**.
- All 10 dialler unit tests pass.

---

## What each box needs

- **Boxes that use the dialler:** update to **v1.1.185+**. The worker keeps full
  functionality (it opts into events); short-lived consumers no longer leak.
- **Boxes that do NOT use the dialler (e.g. meyers):** update to v1.1.185+ as
  well. `pbx-update.sh` will automatically **stop and disable** the dialler
  worker service when the module is off — which removes the only remaining
  events-on connection entirely. After updating, do one Asterisk restart at 0
  active calls to clear the pre-fix orphaned subscriptions.

---

## Deploy + verify steps

```bash
cd /var/www/html
sudo ./scripts/pbx-update.sh          # pulls code, stops/disables worker if dialler off, restarts queue worker
sudo systemctl restart apache2        # mod_php: clears opcache so HTTP paths load new code

# Restart Asterisk once, at 0 active calls, to clear pre-fix orphans:
asterisk -rx "core show channels count"   # proceed only if 0 active
asterisk -rx "core restart when convenient"
```

---

## Monitoring

```bash
# Leaked subscription count — should stay flat
# (0, or 1 while the dialler worker holds its single connection)
asterisk -rx "core show taskprocessors" | grep -c "stasis/p:manager:core"

# Pool-control queue depth — should stay low at idle, grow only with live call volume
asterisk -rx "core show taskprocessors" | grep "stasis/pool-control"
```

There is also an early-warning email (`pbx:asterisk-health-alert`, every 5 min)
that alerts before the pool-control queue reaches the 500 cap.

---

## Commit reference

| Commit | Tag | Change |
|---|---|---|
| `cd48268` | v1.1.182 | `Events: off` on `Asterisk\AmiClient` |
| `c5c641b` | **v1.1.185** | **dialler AMI events opt-in — the definitive third-path fix** |
| `28c07ff` | v1.1.186 | restart queue workers on deploy so new code actually loads |

---

## Files touched by the final fix

- `app/Services/Dialer/AsteriskAmiService.php` — `Events: off` default + `subscribeToEvents()`
- `app/Services/Dialer/Contracts/AsteriskAmiInterface.php` — added `subscribeToEvents()` to contract
- `app/Services/Dialer/FakeAmiService.php` — no-op `subscribeToEvents()` for tests
- `app/Console/Commands/DialerWorkerCommand.php` — worker calls `subscribeToEvents(true)` before connect
- `scripts/pbx-update.sh` — restart queue workers on deploy
