$ oc get nodes -l kubevirt.io/schedulable=trueTroubleshooting the Secure Agent Workspace pattern
Use this page to diagnose common issues when you deploy or use this pattern. Most commands take the user’s name as OPENSHELL_SAW_NAME; the examples use alice.
The workspace VM does not start
- Symptom
oc get vm -n saw-aliceshowsErrorUnschedulableor stays inScheduling, or the OpenShift Virtualization Operator reports that no node supports virtualization.- Cause
The worker nodes are not bare metal, or hardware virtualization is turned off in the firmware. OpenShift Virtualization cannot run the VMs on virtualized nodes.
- Resolution
Add bare-metal worker nodes (on public clouds,
*.metalinstance types) and check that the nodes are labeled as able to run VMs:For more information, see Cluster sizing.
The installer in the VM fails
- Symptom
make openshell-saw-status OPENSHELL_SAW_NAME=aliceshows theinstallorapplystep in a phase other thanDone. The command prints each step with its phase and the first line of the error.- Cause
Common causes are the following:
A provider key is missing. The workspace needs the
inferenceandweb-searchSecrets insaw-alice.The provider in the Secret does not match the profile. The installer refuses an
inferenceSecret whoseproviderdoes not match the profile’s provider type.An image cannot be pulled. The VM pulls the OpenShell images from quay.io and the sandbox images from their registries.
- Resolution
Read the installer’s log to find the cause:
$ make openshell-saw-logs OPENSHELL_SAW_NAME=aliceFor a shell on the VM, run
make openshell-saw-vm-ssh OPENSHELL_SAW_NAME=alice. It adds your SSH key to the VM when needed.For a missing key, check that the ExternalSecrets synced (
oc get externalsecret -n saw-alice) and that the keys are in Vault. After you change yourvalues-secretfile, run./pattern.sh make load-secrets.For a provider mismatch, set the
providerof theinferencesecret in yourvalues-secretfile to the provider type of the profile, and load the secrets again.For an image that cannot be pulled, check that the cluster can reach the registry.
After you fix the cause, restart the VM to run the installer again:
$ make openshell-saw-restart OPENSHELL_SAW_NAME=alice
A sandbox stays in Error
- Symptom
openshell sandbox listshows a sandbox inError, and its supervisor log reportsStartup configuration did not stabilize after 5 attempts.- Cause
With OpenShell 0.1.x, a provider profile that carries more than one annotation makes the gateway compute a different provider revision on every request (NVIDIA/OpenShell#3929). The governance interceptor that this pattern builds keeps only one annotation per profile. An interceptor built without that change, or from another OpenShell release, causes this error.
- Resolution
Check that the governance interceptor runs the image defined in the chart, and that the image and gateways are built using the same OpenShell release:
$ oc get deploy governance-interceptor -n openshell-agents -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'Then restart the VM. The installer recreates sandboxes that are in
Error.
Sandboxes and providers cannot be created
- Symptom
openshell sandbox createoropenshell provider createfails, or the installer logs that the gateway has no profile for a provider type.- Cause
The governance interceptor is unreachable or does not serve that profile. The gateways fail closed.
- Resolution
Check the interceptor pod and its log. The log lists the profiles it loaded:
$ oc get pods -n openshell-agents -l app.kubernetes.io/name=governance-interceptor $ oc logs -n openshell-agents -l app.kubernetes.io/name=governance-interceptor --tail=20If the pod is in
ImagePullBackOff, check that the interceptor image exists and that the cluster can pull it without credentials. A provider type that has no profile incharts/governance-policy/profiles/cannot be created.
An agent cannot reach a model or tool API
- Symptom
The agent reports a network error or
403, or the sandbox log shows a denied connection.- Cause
Sandboxes can reach only the hosts that their attached providers' profiles list, and only from the programs that the profiles list.
- Resolution
Check that the provider is attached to the sandbox (
openshell sandbox provider list <sandbox>). Check also that its profile lists the host and port and the real path of the program that connects, such as/usr/bin/node-*for OpenClaw. For a custom model endpoint, theopenaiprofile must name the endpoint’s host. See Ideas for customization.
The openshell CLI reports a version mismatch
- Symptom
The
openshellCLI cannot connect to the gateway.- Cause
The CLI is from a different OpenShell release series than the gateway. The gateway that this pattern installs is from the 0.1.x series.
- Resolution
Check the CLI version with
openshell --version, and install a 0.1.x CLI from the OpenShell releases.
The openshell CLI reports an expired sign-in
- Symptom
The
openshellCLI reports an expired or inactive token.- Cause
The Keycloak session that the CLI signed in with has ended.
- Resolution
Sign in again:
$ openshell gateway login alice
The openshell CLI reports a TLS handshake error
- Symptom
The
openshellCLI fails with a TLS handshake error.- Cause
The gateway was registered with an old route.
- Resolution
Register the gateway again. The route is
alice-gatewayinsaw-alice.$ make openshell-saw-configure-gateway OPENSHELL_SAW_NAME=alice
The web UI is blank after an upgrade
- Symptom
The OpenShell web UI shows a blank page after the pattern is upgraded.
- Cause
The browser is using cached files from the previous build.
- Resolution
Reload the page without the cache (for example, Shift+Reload), or clear the site data.
The web UI reports that a workspace is not found
- Symptom
The OpenShell web UI shows
workspace ' <name>' not found.- Cause
The web UI build and the gateway release do not match. The OpenShell web UI and the gateway must use the same API version.
- Resolution
Upgrade the gateway, or pin a web UI image that matches it.
Argo CD applications are not healthy
- Symptom
./pattern.sh make argo-healthcheckreports applications that are notSyncedandHealthy.- Cause
An application failed to sync or one of its resources is not ready. Each user has three applications,
saw-<user>-secrets,saw-<user>-bom, andsaw-<user>. The shared ones areopenshift-cnv,openshell-keycloak,vault,openshift-external-secrets,governance-policy,governance-interceptor, andsaw-users.- Resolution
Open the application in the Argo CD console and check the sync status and the resources that are not healthy. For a user’s applications, also check the sections above about the VM and its installer.
