Troubleshooting
← Back to Index | Previous: Network and firewall | Next: Appendix →
10.1 Installer logs
| Product | Typical log locations |
|---|---|
| Deslicer AI | Path printed by the installer (under /opt/deslicer/) |
| DAP | /opt/deslicer/var/logs/dap/install.log, /opt/deslicer/var/logs/dap/ansible.log |
Collect timestamps and the failing step name before contacting support. Prefer installer --status / --troubleshoot over ad-hoc host commands (Chapter 5 §5.6).
Installer self-update: Permission denied on /opt/deslicer/bin
Symptoms: replacing /opt/deslicer/bin/deslicer-dai-install.sh with published installer then mv: … Permission denied, and the run continues on an old version (for example 1.4.8).
Cause: /opt/deslicer/bin is not writable by the admin (missing Host prep, or group membership not active in the current shell). Older DAI installers treated a failed mv as success and re-executed the stale script.
Fix: Complete Host prep so /opt/deslicer/bin is 2770 deslicer:deslicer and you are in group deslicer. Re-login or newgrp deslicer, verify test -w /opt/deslicer/bin, then re-download with inode replace (curl -o *.new, chmod 775, mv -f). Do not rely on /tmp as the long-lived pin. Current installers fail closed if they cannot replace the on-disk copy. See § /opt/deslicer layout below.
/opt/deslicer layout (CIS)
Symptoms: Permission denied on install scripts, curl cannot write bin/, installer dies naming an expected path/owner/mode, or sudo -u deslicer cannot read etc/age/keys.txt.
Expected state (fail-closed if missing):
| Path | Mode | Owner |
|---|---|---|
/opt/deslicer, bin/, var/logs/{dap,ai,ansible,dai}/, var/lib/{dap,dai}-install/, dai-install/ | 2770 | deslicer:deslicer |
/opt/deslicer/etc, etc/age, etc/dai | 0700 | deslicer |
etc/age/keys.txt, etc/dai/provision.enc.yml, ai/install-state.enc.yml | 0600 | deslicer |
ai/, dap/ compose | 0750 | deslicer |
The invoking admin must be in group deslicer (id -nG | grep -qw deslicer) and test -w /opt/deslicer/bin. sudo does not replace group membership.
Cause: CIS / STIG umask 027 leaves new dirs 0750. Product assumes Host prep already created the layout. Installers do not chown -R, reclaim workspaces, or fall back to $HOME/.config/age.
Fix: Re-run greenfield Host prep from Control only on an empty tree. If the host already has an old admin-owned layout, contact Deslicer Support — do not copy keys from $HOME or chown -R /opt/deslicer.
Then re-login SSH (or newgrp deslicer) and verify:
id -nG | grep -qw deslicer && test -w /opt/deslicer/bin && echo OK
stat -c '%a %U:%G' /opt/deslicer /opt/deslicer/bin
sudo -u deslicer test -r /opt/deslicer/etc/age/keys.txt
Download and run installers without sudo as Control prints.
CIS smoke checklist (Host prep + no-sudo pin)
On a CIS / STIG host (umask 027), after Host prep and a fresh SSH session:
id -nG | grep -qw deslicer
test -w /opt/deslicer/bin
stat -c '%a %U:%G' /opt/deslicer /opt/deslicer/bin # expect 2770 deslicer:deslicer
stat -c '%a %U:%G' /opt/deslicer/etc /opt/deslicer/etc/age
# expect 0700 deslicer:deslicer
curl -fsSL https://artifact-registry.deslicer.io/install/linux/enterprise/deslicer-dap-install.sh \
-o /opt/deslicer/bin/deslicer-dap-install.sh.new
chmod 775 /opt/deslicer/bin/deslicer-dap-install.sh.new
mv -f /opt/deslicer/bin/deslicer-dap-install.sh.new \
/opt/deslicer/bin/deslicer-dap-install.sh
bash /opt/deslicer/bin/deslicer-dap-install.sh --version
All steps after Host prep must succeed without wrapping the installer in sudo. If test -w fails, re-login or newgrp deslicer before continuing.
Ansible Galaxy: Permission denied on /etc/ansible
Symptoms: Installer reaches ensuring ansible collections under /opt/deslicer/var/lib/{dap,dai}-install/.ansible/collections, then:
[WARNING]: Skipping Galaxy server https://galaxy.ansible.com. … Permission denied: '/etc/ansible'
[ERROR]: Unknown error when attempting to call Galaxy at 'https://galaxy.ansible.com/api/': [Errno 13] Permission denied: '/etc/ansible'
… ERROR: failed to install required ansible collections
Cause (two layers):
- Without path pins, Galaxy may try to create its cache under
/etc/ansible. - Even with cache pinned under
/opt/deslicer/…, Ansible’s TLS CA discovery always callsos.listdir('/etc/ansible')when that directory exists. On CIS hosts/etc/ansibleis often root-only;listdirraisesPermission deniedand Galaxy wraps it as a failed call tohttps://galaxy.ansible.com/api/.
Fix: Use DAI installer ≥ 1.4.36 and DAP installer ≥ 1.9.47. Those versions pin Ansible home/cache under /opt/deslicer/var/lib/{dai,dap}-install/.ansible and install a venv sitecustomize that skips unreadable CA-search directories (so root-only /etc/ansible no longer breaks Galaxy). Re-download the installer to /opt/deslicer/bin (no sudo) and re-run. Do not chmod or create world-writable /etc/ansible as a workaround.
Confirm after refresh:
bash /opt/deslicer/bin/deslicer-dap-install.sh --version # expect ≥ 1.9.47
bash /opt/deslicer/bin/deslicer-dai-install.sh --version # expect ≥ 1.4.36
ls -ld /etc/ansible 2>/dev/null || true # often exists root-only on CIS
10.2 Temporary-file permission failures
Symptoms: Install fails while escalating to the service account with a message about temporary files and chmod: invalid operator.
Cause: The host is missing setfacl (package acl).
Expected path: Current DAI/DAP installers install acl during host prep. Re-download the published installer and re-run.
If preflight cannot install the package (air-gapped / locked repos), install it once with sudo, then re-run:
# Ubuntu / Debian
sudo apt-get update && sudo apt-get install -y acl
# RHEL / Rocky / Alma / Oracle / Amazon Linux 2023
sudo dnf install -y acl
command -v setfacl
10.3 docker.service fails to start on RHEL
Symptoms: DAI (or DAP) install fails during Docker host prep with Job for docker.service failed / dockerd is not responding, typically on RHEL 10 / Rocky / Alma / Oracle after Docker CE packages are installed.
Cause: Docker’s networking stack needs kernel modules (br_netfilter, overlay, xt_addrtype, iptable_nat) and iptables-nft. Fresh RHEL / AWS images sometimes lack a loadable xt_addrtype (often from kernel-modules-extra, or kernel-uek-modules-extra on Oracle UEK kernels), or the host still runs an older kernel after a newer kernel RPM was installed. Oracle Linux with UEK may already have addrtype built in — installing the RHCK kernel-modules-extra package is unnecessary there.
Typical journal line (kernels may already match):
iptables ... -m addrtype --dst-type LOCAL -j DOCKER: Warning: Extension addrtype revision 0 not supported, missing kernel module?
Expected path: The installer installs Docker and required networking support. If Docker still fails, use the fix below, or reboot when the installer shows REBOOT REQUIRED, then re-run the same enroll/install command. Docker repair is re-running the installer—not hand-pasting Docker install commands from Control.
Fix when journal mentions addrtype / xt_addrtype (even if uname -r matches the newest kernel):
uname -r
find /lib/modules/"$(uname -r)" -name 'xt_addrtype.ko*'
# Typical RHEL / Rocky / Alma / Oracle:
sudo dnf reinstall -y "kernel-modules-extra-$(uname -r)" || sudo dnf reinstall -y kernel-modules-extra
# Oracle / *uek* in uname -r:
# sudo dnf reinstall -y "kernel-uek-modules-extra-$(uname -r)" || sudo dnf reinstall -y kernel-uek-modules-extra
sudo depmod -a
sudo modprobe xt_addrtype
sudo systemctl reset-failed docker
sudo systemctl enable --now docker
sudo docker info
After the installer stops for reboot (newest installed kernel ≠ uname -r):
sudo reboot
# Then re-run the same enroll/install command you used before the reboot.
Ubuntu / Debian / Amazon Linux 2023: this failure mode is uncommon; keep using the existing distro Docker paths (no reboot for modules is usually required).
10.4 Image pull / registry auth
Symptoms: unauthorized, denied, or pull timeouts during install.
Checks:
- Confirm outbound HTTPS to
container-registry.deslicer.io(Chapter 9) - Confirm container registry username/password in the provision package match what Deslicer issued (contact Deslicer to re-export the package if wrong)
- Re-run install only after fixing credentials—do not paste passwords into chat
10.5 Age decrypt failures
Symptoms: DAI or DAP installer cannot decrypt provision.enc.yml, or DAI cannot read install-state.
Checks:
- Age identity is
/opt/deslicer/etc/age/keys.txt(0600 deslicer) and matches the public key used to encrypt the file - For DAP: you are decrypting the Control-issued bundle for this enroll, not an old handoff file
- For DAI:
/opt/deslicer/etc/dai/provision.enc.ymland/opt/deslicer/ai/install-state.enc.ymlare0600 deslicer;--updateuses the same identity sudo -u deslicer test -r /opt/deslicer/etc/age/keys.txtsucceeds. The invoking admin does not readkeys.txtdirectly
10.6 Control or app not reachable on HTTPS
Symptoms: Browser errors on https://<dai-host>/ or /control, while the host-local health check works.
Checks:
- On the DAI host:
bash /opt/deslicer/bin/deslicer-dai-install.sh --status - Confirm DNS points at this host
- Confirm certificate mode in Control (ACME needs port 80; PEM needs valid files under
/etc/caddy/certs/<hostname>/— see Chapter 6 §6.3) - Re-apply Web certificates using the Control-generated host command (Chapter 6)
10.7 ACME failures
Symptoms: Automatic HTTPS never becomes ready; Control or the installer report certificate challenge failures.
Checks:
- Port 80 reachable from the public Internet to this host
- No other process bound to 80/443
- Hostname in install answers matches the DNS name you challenge
10.8 DAP enroll probe failures
Symptoms: deslicer-dap-install.sh cannot reach --dai-url.
Checks:
- DAI Control is up on HTTPS
- DAP host can resolve and reach the DAI hostname (firewall between hosts)
- Enroll token has not expired—mint a new token in Control
10.9 Fresh token vs existing volumes
Symptoms: Fresh install refuses because DAP volumes already exist, or Postgres auth fails after a rotated Fresh token.
Checks:
- Prefer an Update token from Control when
/opt/deslicer/dapalready has data volumes - If a previous Fresh run left volumes and host config, and you re-run the same Fresh token whose Postgres password still matches that config, the installer resumes (preserves secrets, finishes bootstrap, and brings the stack up if needed)
- Fresh that rotates secrets against a live volume whose password differs will refuse—wipe volumes for a clean Fresh, or use Update/Repair
- Do not re-use a Fresh token that rotated the DB password against a live stack that still has the previous password
Encrypted secret verify failures: If verify reports a missing portal row or keyring mismatch, upgrade the installer / DAP install package and contact support with logs (§10.21).
10.10 Update token cannot read compose .env
Symptoms: --update fails with missing or unreadable /opt/deslicer/dap/.env (often Permission denied for a non-root user).
Checks:
- Run the Control-printed install command as the sudo-capable admin
- Confirm
.envexists under/opt/deslicer/dapwith mode0600and is readable by the install path Control documents - Re-download the installer from Control if the host copy is older than the pin Control shows (
deslicer-dap-install.sh --version)
10.11 DNS / dns_error vs Healthy
Symptoms: Control shows dns_error, failed, or Inactive while Retest shows Healthy.
Checks:
- Public Observer hostname resolves from the DAI host to the DAP host IP
- On split installs, Deslicer AI probes the public Observer HTTPS URL Control shows—not a host-local name
- After DNS/TLS is fixed, use Retest connection or Retry probe so Control sets install status to active and clears the stale error
- If ACME backed off while DNS was wrong, restart Caddy on the DAP host after DNS is correct
10.12 Deslicer AI reports DAP down
Symptoms: Control or dashboard shows Observer unreachable.
Checks:
- Open the Observer health URL Control shows for the backend
- Confirm the backend URL in Control matches a URL the DAI host can reach (split: public edge HTTPS on 443)
- Confirm certificate trust on the DAI host for that Observer URL
- On each host, confirm the Compose project is running (
docker composeunder/opt/deslicer/aiand/opt/deslicer/dap) - Run Retry probe / Retest after fixing network or TLS so install status can become Active
10.13 Login returns 403 {"error":"Forbidden"}
Symptoms: Browser sign-in (email/password local login on https://<dai-host>/ or /control) fails immediately with:
{"error":"Forbidden"}
HTTP status 403. Wrong password returns 401 (Invalid credentials) instead—this section is only for 403 Forbidden.
Cause: Local login accepts the browser origin only when it exactly matches the public app URL configured for the Deslicer AI stack (no trailing slash). Typical mismatch after a hostname rename, DNS fix, or certificate-only apply that updated Caddy but not the application configuration.
Checks:
- Note the URL in the address bar (scheme + host only), for example
https://dai.example.com - On the DAI host, compare the Caddy site name and the app URL:
# Public edge hostname (what Caddy serves)
sudo awk '/^[A-Za-z0-9][A-Za-z0-9._-]*[[:space:]]*\{/{ gsub(/\{/,"",$1); print $1; exit }' /etc/caddy/Caddyfile
# App-allowed origin (must match the browser URL)
sudo grep '^NEXT_PUBLIC_APP_URL=' /opt/deslicer/ai/.env
-
Confirm
provision.yml/provision.enc.ymlhas the same public hostname in both:deployment.urls.daiPublic(preferred; drives compile / Caddy)app.domainName(keep in sync for handoffs)
-
Optional origin probe (expect 401
Invalid credentialswhen Origin matches; 403 when it does not):
curl -skS -X POST "https://<dai-host>/api/auth/local-db/login" \
-H 'content-type: application/json' \
-H 'origin: https://<dai-host>' \
-d '{"email":"probe@example.com","password":"wrong"}'
Fix:
- Correct
daiPublic/domainNamein the provision file (exact DNS name, including spelling). Contact Deslicer to re-export the package if you cannot edit the handoff safely. - Re-run a full DAI update so application configuration and containers pick up the new URL—omit
--proxy-only:
bash /opt/deslicer/bin/deslicer-dai-install.sh --update
Confirm the installer log ends with a full install completion message, not only a certificate-only apply.
- Verify
NEXT_PUBLIC_APP_URLmatches the browser URL, then sign in again (hard-refresh if the tab still posts from an old origin).
Related failures (not this 403):
| Response | Meaning |
|---|---|
404 {"error":"Not available"} | Local login is not enabled for this engagement |
401 {"error":"Invalid credentials"} | Wrong email/password (or user not seeded) |
429 account locked | Too many failed attempts—wait and retry |
| Identity provider errors | Separate path when your engagement uses external SSO—check IdP redirect URIs against the public DAI URL |
10.14 Control DAP page returns 500 (install-package cache)
Symptoms: https://<dai-host>/control/platform-integrations/dap or Updates → DAP fails with HTTP 500. Public application health may still be healthy.
Cause: The Control install-package cache directory is not writable by the Control service.
Checks / fix:
bash /opt/deslicer/bin/deslicer-dai-install.sh --status
bash /opt/deslicer/bin/deslicer-dai-install.sh --troubleshoot
bash /opt/deslicer/bin/deslicer-dai-install.sh --troubleshoot --fix-dap-bundle-cache
Then reload the Control DAP / Updates page.
10.15 Interactive schema: create vs rename
Symptoms: During DAI --update, Drizzle asks:
Is <table> table created or renamed from another table?
Action: Choose create table. Do not choose rename unless Deslicer Support told you to for that specific table and release.
Why: Rename guesses are often wrong (unrelated tables that happen to be missing vs new). Rename can drop or merge production data.
If the wrong choice was already confirmed: stop and contact support before re-running; do not invent a second rename to “undo” it.
10.16 DAP install package Apply / serve failures
Symptoms: Control shows Update available but Apply fails; DAP install fails on missing or mismatched checksum; enroll downloads return 503.
Checks:
- Apply update requires outbound HTTPS from the DAI host to
registry.deslicer.ioandartifact-registry.deslicer.io(Chapter 9). Air-gapped sites use Import package under Control → Updates → DAP instead. - If Apply says the package needs a newer Deslicer AI, update the DAI image first—Control does not skip that gate.
- Applying a release that is not the registry latest requires explicit confirmation in Control.
- Messages about a missing package checksum, or HTTP 503 indicating no package is staged, mean Control has no verified package. Fix with
bash /opt/deslicer/bin/deslicer-dai-install.sh --repair, or Import package, then re-issue the enroll token. - Package download 503 with an integrity/cache message: the on-disk cache no longer matches the stored package. Prefer Roll back, then Apply update again.
- After Apply or repair, confirm Control → Updates → DAP shows a staged checksum before re-running DAP enroll /
--update. - If Control itself 500s before you can open Updates, fix cache permissions first (§10.14).
Routine DAP package updates do not require a new Deslicer AI image. Schema or enroll API changes still do.
10.17 No match for argument: python3.13 on RHEL 10 (AWS)
Symptoms: Host prep fails with:
No match for argument: python3.13
No match for argument: python3.13-pip
Error: Unable to find a match: python3.13 python3.13-pip
Often paired with Unable to read consumer identity on AWS RHEL images (RHUI still serves BaseOS/AppStream).
Cause: RHEL 10 BaseOS/AppStream ships Python 3.12. python3.13 needs additional repositories (CodeReady Builder / CRB and EPEL). A bare sudo dnf install -y python3.13 … without those repositories always fails. Age can install successfully while Python install fails.
Fix: do not install Python by hand. Re-download the current install script and re-run the same enroll/install command—the installer enables the required repositories and installs Python 3.13, pip, curl, and acl. If prep still fails, capture the full installer output and the result of:
dnf repolist --enabled
cat /etc/os-release
and see Chapter 3 / §10.21. Control’s age install commands are not a Python repair path.
10.18 Listed or synced model fails in chat
Symptoms: A model row shows synced (or appeared in Discover / Bulk import), Test connection on the provider is green, but chat or the per-model Test action fails with AWS AccessDeniedException, validation errors, or gateway 404.
Cause: Registration and listing do not prove invokability.
| Failure class | What it means | What to check |
|---|---|---|
| Entitlement | Account has not been granted that model in the connection region | Commercial: Marketplace permissions; wait ~15 minutes after first invoke. GovCloud: Bedrock console → Model access. Do not send commercial operators to the GovCloud Model access page. |
| IAM | Credential lacks an action (bedrock:InvokeModel, streaming, or gateway auth) | Fix IAM policy; Test connection on non-gateway kinds does not call the provider — use per-model Test |
| Wrong upstream id | Bare foundation-model id where AWS expects a cross-region inference profile (bedrock/eu.…) | Use the profile id AWS documents; see Chapter 13 §13.5 |
| Custom gateway prefix | Model ID and Upstream disagree, or api_base omitted /v1 | Same gateway id in both fields; base URL ends with /v1; Sync now after product update |
Fix: On Enterprise → AI providers & models, run Sync now, then Test the row (not Test connection alone). See Chapter 8 §8.5 and Chapter 13 §13.7.
10.20 Enterprise support scripts
Deslicer AI ships host-side diagnostics under scripts/enterprise/packaged/ on the DAI install root (typically /opt/deslicer/ai). Run every command below from that directory as a user who can invoke Docker.
Compose rule on provisioned hosts: use docker compose --env-file .env only. Do not pass ad-hoc -f docker-compose.yml lists — that drops the COMPOSE_FILE overlay chain (including deslicer_internal networking with DAP). The scripts follow this rule internally.
| Script | When to use | Example invocation |
|---|---|---|
collect-support-bundle.sh | Before opening a support ticket; stack unhealthy; you need compose state and service logs for Deslicer Support | bash scripts/enterprise/packaged/collect-support-bundle.sh |
diagnose-litellm-model.sh | Per-model Test or chat fails after Sync now; custom gateway routing; you need LiteLLM /model/info plus a real completion probe | bash scripts/enterprise/packaged/diagnose-litellm-model.sh <model_id> |
verify-gateway-credentials.sql | Custom gateway rows show synced but completions fail; after credential rotation; sync_status is failed or drifted | See Custom gateway SQL check below |
Support bundle (collect-support-bundle.sh)
Collects compose snapshots, per-service logs for the full DAI stack (and DAP when /opt/deslicer/dap is present), container network attachments, and redacted env files. The collector auto-detects a provisioned host (.env + overlay chain) or falls back to a repo checkout (docker-compose.enterprise.yml + .env.enterprise).
cd /opt/deslicer/ai
bash scripts/enterprise/packaged/collect-support-bundle.sh
bash scripts/enterprise/packaged/collect-support-bundle.sh --output /tmp/deslicer-support.tar.gz
bash scripts/enterprise/packaged/collect-support-bundle.sh --log-lines all
Redaction: support-bundle-redaction.sh strips API keys, tokens, passwords, and private keys by key name and by literal match against env values. The archive is refused if any known secret survives. You do not need to redact the tarball manually before sending it to support — still avoid attaching raw .env files alongside the bundle.
Attach the generated .tar.gz when contacting support (§10.21).
LiteLLM model probe (diagnose-litellm-model.sh)
Mirrors what the Enterprise workspace Test action checks at the proxy: GET /model/info (deployment summary: model_name, upstream litellm_params.model, api_base) then POST /v1/chat/completions with max_tokens=5.
- Probes via
nodeinsideweb→http://litellm:4000(same path as DAI chat), orpython3insidelitellmwhenwebis down. - Does not require
curlin containers. <model_id>is the Model ID / LiteLLMmodel_nameregistered by Sync now (for examplegpt-4o-mini,openai/gpt-5-nano, or a catalog slug). Omit the argument to probe the first chat-capable deployment in/model/info.
cd /opt/deslicer/ai
bash scripts/enterprise/packaged/diagnose-litellm-model.sh
bash scripts/enterprise/packaged/diagnose-litellm-model.sh bedrock-claude-4-6-sonnet
bash scripts/enterprise/packaged/diagnose-litellm-model.sh openai/gpt-5-nano
For custom gateway prefix and api_base /v1 mistakes, see Chapter 13 §13.4 and §10.18.
Custom gateway SQL check
Read-only Postgres checks for custom gateway connections: secret_config_key pointer, encrypted system_configs row metadata, and catalog sync_status. Does not decrypt API keys — use Enterprise → AI providers & models → Test connection → Sync now → Test to validate decrypted credentials.
cd /opt/deslicer/ai
docker compose --env-file .env exec -T app-db \
psql -U deslicer -d deslicer -v ON_ERROR_STOP=1 -f - \
< scripts/enterprise/packaged/verify-gateway-credentials.sql
If your install uses non-default DB user/database names, substitute the values from /opt/deslicer/ai/.env (APP_DB_USER, APP_DB_NAME).
10.21 Getting help
Contact support@deslicer.com with:
- Engagement / hostname
- Installer versions (
deslicer-dai-install.sh/deslicer-dap-install.sh --version) and output of--status/--troubleshootwhen relevant - Support bundle from §10.20 when the stack or LLM routing is involved
- Applied DAP install-package version from Control (Updates), if relevant
- Redacted log excerpts (no passwords, tokens, or private keys)
- For login 403: browser URL and the configured public app URL (no secrets)
← Back to Index | Previous: Network and firewall | Next: Appendix →