# IMPORTANT FIXES — Deferred / To Do

A running list of known issues and improvements that are **parked** (agreed but not
yet implemented). Each entry says what the problem is, why it matters, where it
lives, and the proposed fix. Newest at the top.

---

## RESOLVED — Long call recordings would not transcribe (fixed v1.1.17)

**Status:** DONE.

**Symptom:** short calls transcribed fine, but a ~7-minute recording never
transcribed no matter how many times it was retried.

**Root cause:** the whisper command was hard-killed by a FIXED `timeout 300`
(5 min) while running single-threaded at lowest priority (`--threads 1` +
nice/ionice). A 7-minute recording takes longer than 5 minutes to transcribe,
so `timeout` killed it before any JSON was produced → "FAILED, no JSON", and
every retry hit the same wall. Short calls finished within 5 minutes, which is
why they worked.

**Fix (v1.1.17):** in `app/Services/TranscriptionService.php`
- The timeout now SCALES with the recording length: effective timeout =
  `max(WHISPER_TIMEOUT, audio_seconds * WHISPER_TIME_FACTOR + 60)`, capped at
  `WHISPER_MAX_TIMEOUT` (default 1800s). Duration is read from the WAV header.
- Thread count is configurable (`WHISPER_THREADS`, default 4) instead of a
  hardcoded 1, so long recordings finish in reasonable time. Still `nice -n 19`
  + `ionice -c3`, so live calls always win the CPU (the original freeze guard —
  per-file lock, concurrency cap, load guard — is untouched).
- New env keys added to `config/services.php`, `.env.example`, and
  `pbx:ensure-env`. Works via config defaults even without `.env` entries.
- START log now records detected duration, chosen timeout, and threads.

---

## LAN-mode deployment profile (desk-phones-only, no domain / no HTTPS)

**Status:** DONE (v1.1.27) — `scripts/pbx-setup-client.sh --lan-mode`. The setup
script now has a LAN path that:
- prompts for the **LAN IP** (not a hostname) and an optional **WAN IP** for the
  SIP trunk's NAT;
- writes `.env` with `APP_URL=http://<lan-ip>`, `SESSION_SECURE_COOKIE=false`
  (so login works over HTTP), and an empty `WEBRTC_WSS_URL`;
- writes an **HTTP-only Apache vhost** (no HTTPS redirect, no WSS proxy);
- **skips** Let's Encrypt issuance and the HTTPS vhost entirely;
- points PJSIP `external_*_address` at the WAN IP when given (else the LAN IP)
  and `local_net` at the LAN subnet;
- derives the machine hostname from the company name (not the IP);
- prints LAN-specific notes (webphone/mobile unavailable; how to add later).
The browser webphone and mobile app are intentionally unavailable in this mode
(they need HTTPS + a cert). Re-running setup WITHOUT `--lan-mode` after assigning
a domain + cert promotes the site to full mode. `pbx:setup-mobile-transport`
already self-skips when no cert is present, so updates are safe on a LAN box.

**Original rationale / what was needed (kept for context):** Needed for clients who
run ONLY desk phones on their local network (no webphones, no mobile app), with
the PBX on a private LAN IP and outbound internet access for the SIP trunk.

### What ALREADY works on a local IP / HTTP (verified in code)
- Desk phones over SIP UDP/TCP (5060) — no cert/domain needed.
- Web GUI over plain HTTP — there is NO forced-HTTPS in the app code (no
  forceScheme, no HTTPS-redirect middleware) and `SESSION_SECURE_COOKIE=false`
  by default, so login/sessions work over `http://<lan-ip>`.
- Outbound/inbound calls via the SIP trunk — needs only outbound internet, not
  an inbound domain/public IP.
- Recording, voicemail-to-email, CDR, dialplan, PIN, blacklist — all internal.
- CNAM caller-name lookup — uses `http://127.0.0.1/api/cnam`, domain-independent.

### What does NOT work without a domain + cert (by design — fine for this profile)
- Browser webphone (WebRTC over WSS) — browsers refuse WSS without HTTPS+cert.
- Linphone mobile transport (`transport-tls-mobile`) — bound to a cert.
  → Acceptable for a desk-phones-only client; just confirm they will never want
    a webphone/mobile later (adding it means adding a domain + cert).

### Things to formalise for the LAN profile
1. **`APP_URL` must be the LAN IP** (e.g. `http://192.168.1.50`), not a domain —
   otherwise GUI-generated absolute URLs/email links point somewhere unreachable.
2. **Setup script needs a `--lan-mode` path** that SKIPS: Let's Encrypt issuance,
   hostname collection, and the mobile TLS transport (`pbx:setup-mobile-transport`
   already warns + skips when no cert is found, but the overall flow still assumes
   a domain). Proposed: a flag that writes `APP_URL=http://<ip>` and bypasses the
   cert/hostname steps.
3. **Confirm the SIP trunk works through the client's NAT** (inbound + outbound)
   — same consideration as any on-prem PBX. `rtp_symmetric` / `rewrite_contact`
   are already set on trunks. Desk-to-desk calls stay on the LAN (no NAT).
4. **Confirm outbound SMTP** works from the client LAN (voicemail emails).
5. **Decide remote admin access** for us (the provider): VPN to the LAN, or
   accept LAN-only GUI access. No domain = no remote GUI without a VPN.
6. GUI SSL / Let's Encrypt settings tabs will be unused — note this for the client.

### Acceptance
- A fresh box on a private IP with no domain: desk phones register and call
  internally + via trunk; GUI reachable over HTTP on the LAN; voicemail emails
  send; no failed cert/hostname steps during install.

---

## High-concurrency (100+ extensions) — status after v1.0.90

A concurrency audit was done ahead of a 125-extension client. Status of each item:

**DONE in code (auto-applied on deploy via `pbx:tune-performance`):**
- PJSIP threadpool (initial 20 / max 150 / inc 10) — written into the existing
  `[system]` block of `pjsip.conf`.
- Stasis threadpool (initial 10 / max 100) — `stasis.conf`. Fixes the
  `stasis/pool-control` backlog (CDR/AMI lag, "queue reached 500" log floods).
- ODBC pool raised 30 → 50 — `res_odbc.conf`.
- `maxcalls=200` / `maxload=4.0` — `asterisk.conf [options]`.
- All env-overridable (`PBX_*` in `.env.example`); idempotent; re-run safe.
- `func_audiohook` AUDIOHOOK_INHERIT guard — done v1.0.88 (only emitted if the
  module is loaded, so inbound recording never crashes the call).
- `asterisk_dialplan` compound index — already the PRIMARY key
  `(context, exten, priority)`, so realtime lookups are already optimal. No-op.

**OPERATIONAL — needs an action outside code (do these for the big client):**
1. **RESTART REQUIRED after first deploy.** The threadpool / stasis / maxcalls
   settings load at Asterisk START, not on reload. `pbx:tune-performance` writes
   them and prints "restart required" but deliberately does NOT restart (that
   drops calls). Schedule `sudo systemctl restart asterisk` in a maintenance
   window once, after the v1.0.90 deploy. Verify with
   `asterisk -rx "core show settings"` (Maximum calls should show 200) and
   `asterisk -rx "core show taskprocessors" | grep stasis` (queues near 0).
2. **Recording disk rotation.** ~11 GB/day at 125 ext. DISK ALERTING DONE
   (v1.1.38): `pbx:disk-space-alert` hourly emails at 80/85/90/95% (branded
   template, recipients in Settings → Email Notifications, dedup so each tier
   alerts once).
   **POLICY (customer requirement — do NOT auto-delete):** customers do not want
   automatic pruning/rotation. Recordings are retained until the customer gives
   WRITTEN approval to remove a specific set (e.g. "all recordings 12 months and
   older"). When the disk fills, we contact the customer, get written approval,
   then WE remove the approved set manually. So there must be NO scheduled
   deletion job. The supporting tool (if/when built) should be a MANUAL,
   approval-friendly command: a default DRY-RUN preview (count, total size, date
   range for a chosen cutoff) to attach to the approval request, then an explicit
   `--execute` with typed confirmation that removes audio + `.json` transcript
   files, clears the `cdr.recordingfile` reference / `call_recordings` row so the
   GUI shows no dead play buttons, and writes an audit-log entry (who, when,
   cutoff, count, bytes freed). NO scheduler entry. (Supersedes the earlier
   `recordings:prune` auto-rotation idea.)
3. **RAM.** 7.6 GB is tight for 125 ext with G.729 transcoding. Recommend 16 GB
   for large client deployments (hardware, not code).
4. **`PBX_STASIS_DECLINE_MWI`** is available but defaults OFF. If the big client
   does not use voicemail message-waiting indicators, set it `true` to shed
   stasis load further. Only then — it disables MWI updates.

---

## RESOLVED — Whisper transcription overloaded the PBX (fixed v1.0.85)

**Status:** DONE. Was a production-down issue on the client PBX.

**What happened:** Call-recording transcription launched multiple `whisper`
processes against the SAME .wav from different triggers (5-min cron as root +
web buttons as www-data + queued job), with no locking, no concurrency cap, and
no CPU/IO throttling. Two jobs hit ~280% CPU each and froze the PBX. Whisper was
disabled on the client in an emergency via
`mv /usr/local/bin/whisper /usr/local/bin/whisper.disabled`.

**Fix (v1.0.85):** Rewrote `app/Services/TranscriptionService::transcribe()` with
hard guards, plus a `whisper` config block in `config/services.php` and
documented keys in `.env.example`:
- Master enable flag — `WHISPER_ENABLED` defaults to **false** (opt-in per site).
- Respects the disabled binary (skips if the whisper binary is missing/renamed).
- Atomic per-file lock via `mkdir` — the same recording can never be processed
  twice at once, regardless of which user/trigger launches it. Stale locks past
  the timeout are reclaimed.
- Skip if the output JSON already exists.
- Global concurrency cap (`WHISPER_MAX_CONCURRENT`, default 1) checked before and
  during each batch via `pgrep -fc`.
- `timeout N nice -n 19 ionice -c3 ... --threads 1` — runaway jobs are killed and
  live calls always win CPU/disk.
- Full logging: START / COMPLETE / FAILED / skipped, and the 5-min batch is a
  no-op when disabled or at capacity (batch limit 3, not 10).
- All four triggers (cron, both controllers, queued job) route through this one
  guarded service.

**To re-enable on a client:** set `WHISPER_ENABLED=true` in `.env`, restore the
binary (`mv /usr/local/bin/whisper.disabled /usr/local/bin/whisper`), then
`php artisan config:clear` (or re-cache). Safe now: off by default, one job at a
time, locked against duplicates, throttled, and timeout-guarded.

---

## RESOLVED — Per-extension Call Recording toggle now gates recording (fixed v1.1.18 + v1.1.19)

**Status:** DONE. Recording still records everything by default; turning the
per-extension "Call Recording" toggle OFF now stops recording for that
extension's calls in every direction (inbound, outbound, internal).

**How it was fixed (two steps, both shipped):**
- **v1.1.18 — fixed the argument mangling first.** The full extension now
  reaches `subStartRec` intact in both ARG2 (caller) and ARG3 (callee). The
  cause was `${CALLERID(num)}` parens corrupting the `U()`/Gosub arg parser; we
  capture paren-free vars at priority 1 (`MSet REC_SRC=...,REC_DST=...`) and pass
  `U(subStartRec^INT^${REC_SRC}^${REC_DST})`. Live-traced clean as `(INT,1359,1350)`.
- **v1.1.19 — added the opt-out gate (Option A: "never record this extension").**
  Rather than the ODBC flag (an ODBC read can abort the subroutine), we use
  AstDB family `REC`: `REC/<ext>=0` marks an extension opted OUT. `subStartRec`
  priority 3 sets `REC_OPTOUT=${DB(REC/${ARG2})}${DB(REC/${ARG3})}` and priority
  4 `GotoIf $["${REC_OPTOUT}"!=""]?9` skips MixMonitor when EITHER party opted
  out. A missing/empty key always records — fail-safe, a lookup can never
  silently stop recording. The GUI toggle (`ExtensionFeatureController::toggle`
  / `bulkToggle`) writes/clears `REC/<ext>`, `ExtensionRoutingService::syncRecordOptOuts()`
  reconciles the whole family (batched via AMI), `pbx:deploy` re-syncs it on
  every deploy, and `pbx:audit-recording --fix` rebuilds `subStartRec` to the
  canonical gated form (idempotent, transactional). The seeding migration seeds
  the gated body for fresh installs.
- **Live tested & confirmed:** caller opt-out skipped, callee opt-out skipped,
  neither-opted-out recorded, outbound recorded — all calls connected with audio.

**Behaviour-change note on first deploy:** extensions that already have
`record_calls=0` will STOP recording once v1.1.19 deploys (previously the toggle
did nothing). Review which extensions have recording switched off before/after
the deploy.

---

<details>
<summary>Original parked write-up (kept for context)</summary>

**Severity:** Medium. Recording is functional and reliable today (all directions
record). The gap is that the GUI toggle is misleading — it implies control that
isn't wired up. Needed for clients who must record only some extensions (e.g.
compliance, privacy, or selective-recording requirements).

### What works today (do NOT break this)
- Recording is **unconditional**: every answered call records — inbound,
  outbound, internal — via `U(subStartRec^DIR^caller^callee)` on each Dial.
- Confirmed working with real audio for all three directions (v1.0.80 restored
  the outbound recording handler that had been lost).
- The `subStartRec` subroutine (context in `asterisk_dialplan`) currently records
  with NO condition check — it always starts MixMonitor.

### The problem
- The extension edit screen has a **"Call Recording"** toggle
  (`extension_features.record_calls`, described as "Record all inbound and
  outbound calls for this extension") — but **nothing in the dialplan reads it**.
  Turning it off does NOT stop recording.
- A proper gated subroutine `subRecCheck` exists (seeded by
  `database/migrations/2026_06_03_120000_seed_recording_subroutines_dialplan.php`)
  and a function `ODBC_REC_ENABLED(ext)` exists
  (`/etc/asterisk/func_odbc.conf` → `SELECT COALESCE(record_calls,1) FROM
  extension_features WHERE ext = ARG1`). But **all 47 Dial strings call
  `subStartRec` (unconditional), 0 call `subRecCheck` (gated).**

### Why the obvious fix FAILED (important — do not repeat)
Attempt: gate inside `subStartRec` by looking up `ODBC_REC_ENABLED` for the
caller (ARG2) and callee (ARG3), Return early if neither has recording on.
This **broke recording entirely** because of a PRE-EXISTING ARGUMENT-MANGLING
BUG: the extension number arrives at the subroutine truncated/garbled.

Observed in a live trace for a call to/from 1359:
```
Gosub(subStartRec,s,1(INT,59))      ← should be (INT, 1359, 1359)
  CLEAN_CALLER = 59                 ← should be 1359
  CLEAN_CALLEE = (empty)            ← should be 1359
  REC_CALLER   = (empty)            ← ODBC lookup of "59" finds nothing
  REC_CALLEE   = (empty)
  GotoIf 0?7 → Return (no recording)
```
The old unconditional `subStartRec` "worked" only because it never used those
args for a decision — it just recorded regardless. Any flag-based gate that
trusts ARG2/ARG3 will fail until the mangling is fixed. (Filenames show the same
symptom: recordings are named `...-INT-59-...` instead of `...-INT-1359-...`.)

### Root cause to fix FIRST
The `${CALLERID(num)}` / extension values passed into `U(subStartRec^...)` are
being truncated (1359 → 59) before/at the Gosub. Likely an argument-parsing /
quoting issue in how the `U()` option is built in the Dial appdata, or how
`${CALLERID(num)}` resolves at that point. Until the correct extension reaches
the subroutine, no per-extension lookup can work.

Touch points to investigate:
- `app/Http/Controllers/ExtensionController.php` — `resolveRealtimeDialStep()`
  and `syncHotDeskDialplan()` build `U(subStartRec^INT^${CALLERID(num)}^<ext>)`.
- `app/Services/OutboundRouteDialplan.php` — `U(subStartRec^OUT^${CALLERID(num)}^${DIAL_NUMBER})`.
- `app/Console/Commands/AuditDialplanRecording.php`, `RingGroupController.php`,
  `SpeedDialController.php` — other sites that build the same handler.
- The `subStartRec` / `subRecCheck` bodies are seeded by
  `database/migrations/2026_06_03_120000_seed_recording_subroutines_dialplan.php`.

### Proposed fix (ordered — do them in this order)
1. **Fix the argument mangling** so the full extension (e.g. 1359) reaches the
   subroutine intact in BOTH ARG2 (caller) and ARG3 (callee/destination). Verify
   with a live trace that `Gosub(subStartRec,s,1(INT,1359,1359))` shows the full
   numbers. This also fixes the wrong recording filenames as a bonus.
2. **Only then** gate recording on the per-extension flag. Either:
   (a) add the `ODBC_REC_ENABLED` caller/callee check at the top of
   `subStartRec` (one-place change, instantly covers all 47 call sites), OR
   (b) switch the Dial strings to the existing `subRecCheck` subroutine.
   Option (a) is lower-risk and preferred.
3. **Decide default behaviour for unknown/PSTN parties.** `ODBC_REC_ENABLED`
   returns empty for a non-extension (PSTN number) and for an extension with no
   `extension_features` row. Confirm intended default: record-by-default for a
   known extension (COALESCE already returns 1), and let the OTHER party (the
   real extension) decide on outbound/inbound.

### Acceptance
- Recording still works for inbound, outbound, and internal when the toggle is ON.
- Turning the per-extension "Call Recording" toggle OFF stops recording for that
  extension's calls (verified per direction with real answered calls).
- Recording filenames contain the correct full extension number.
- Deploy-safe: gate lives in a subroutine seeded/maintained by an idempotent
  migration so clients self-heal on update.

### Related history (context)
- v1.0.80 — restored the outbound recording handler (outbound was not recording
  at all). Recording is currently unconditional/always-on across all directions.
- A gating attempt via a migration (`2026_06_05_140000_gate_substartrec_by_record_flag`)
  was tried, broke recording due to the arg-mangling bug above, was rolled back,
  and the migration file deleted. Do NOT reintroduce a gate before fixing the
  arg mangling.

</details>

---

## 1. AstDB routing sync does not scale to large extension counts (200+)

**Status:** RESOLVED — fixed v1.1.16 (batched AstDB writes via AMI).

**What was done:** AstDB writes no longer spawn one `asterisk -rx` process per
key. `ExtensionRoutingService` now has reentrant `beginBatch()`/`flushBatch()`;
all queued `database put/del` commands are sent over a SINGLE AMI connection
(`AmiClient::commandMany()`), each acknowledged before the next so nothing is
dropped. `syncAstDb()`, `syncTrunkStatusToAstDb()`, `syncAllCallerIds()`, the
`pbx:deploy` routing loop, and the `SipTrunkController` trunk-save cascade are
all wrapped in a batch. Measured on dev: 150 dedicated extensions synced in
~1 s (was ~30 s) with 150/150 written. The trunk-save cascade also raises the
time limit for >25 affected extensions as a belt-and-braces guard. Falls back to
per-command `asterisk -rx` if AMI is unavailable.

**Note on transports tried:** piping to `asterisk -r` via stdin was fast but
SILENTLY DROPPED the tail of large batches (134/150) and could hang — do NOT use
it. AMI `Command` responses on Asterisk 22 end with a blank line (`\r\n\r\n`),
not the old `--END COMMAND--` marker.

### Original problem (kept for context)

### What works today
- The outbound digit-strip logic is correct and the double-strip guard is in
  place (see v1.0.79). On a per-call basis there is **no** performance concern —
  the dialplan rows are fixed regardless of extension count.
- On the current small dev box (≈13 extensions, 3 dedicated) everything is fast.

### The problem
`App\Services\ExtensionRoutingService::syncAstDb()` writes each extension's
routing into Asterisk AstDB (family `EXTROUTE`) by shelling out to
`asterisk -rx "database put ..."` — **one process spawn per key**.

- Measured cost: **~16 ms per `asterisk -rx` call**.
- Each extension in `dedicated` mode needs ~6 writes (mode, primary,
  trunk_strip, trunk_prepend, failover keys, failover_enabled) plus one
  `sip_trunks` DB lookup.
- So ~6 × 16 ms ≈ **~100 ms per extension**.

At 200 dedicated extensions that is **~20 seconds** of shell-outs.

### Two places this hurts
1. **Trunk-save cascade** — `App\Http\Controllers\SipTrunkController::update()`
   (added v1.0.77) re-syncs every extension using the trunk **inside the HTTP
   request**. At 200 extensions this can exceed PHP `max_execution_time` (often
   30 s) and **time out / 500**, leaving routing half-applied (some extensions
   updated, others stale). This is the most dangerous one.
2. **Deploy self-heal** — `App\Console\Commands\PbxDeploy::syncExtensionRouting()`
   (added v1.0.79) re-syncs ALL routing rows on every `pbx:deploy`. ~20 s added
   to the deploy. One-off and on CLI (no web timeout), so lower risk, but slow.

### Proposed fix (do both)
1. **Batch the AstDB writes.** Replace per-key `asterisk -rx` calls with a single
   CLI invocation that issues all `database put`/`database del` for an extension
   at once (Asterisk accepts multiple commands per `asterisk -rx`, or pipe them
   via stdin / a temp command file). Expected ~6× speedup → 200 extensions from
   ~20 s to ~3 s. Benefits BOTH the cascade and the deploy.
   - Touch point: `astdbPut()` / `astdbDelete()` / `clearAstDb()` in
     `app/Services/ExtensionRoutingService.php`. Add a batched
     `syncAstDbMany(iterable $rows)` that collects all commands and flushes once.
2. **Offload large trunk-save cascades.** In `SipTrunkController::update()`, when
   the number of affected extensions exceeds a threshold (~25), run the cascade
   via the scheduler/queue instead of inline, OR chunk it and raise the time
   limit for that operation, so a trunk save can never time out mid-cascade.

### Acceptance
- Saving a trunk on a 200-extension client returns quickly and never 500s.
- `pbx:deploy` routing step completes in a few seconds, not tens of seconds.
- All affected extensions end with correct `EXTROUTE` values (no half-applied state).

### Related history (context)
- v1.0.77 — trunk strip change cascades to all extensions using the trunk.
- v1.0.78 — trunk strip/prepend fields show current values in the GUI.
- v1.0.79 — double-strip guard (route vs trunk strip collapsed to one effective
  value, applied once) + `pbx:deploy` self-heal that re-derives every
  extension's `trunk_strip`/`trunk_prepend` from the trunk config.
