Skip to content

Instantly share code, notes, and snippets.

@itachi-re
Created July 26, 2026 20:09
Show Gist options
  • Select an option

  • Save itachi-re/f8f343f5bf0923b8f248d8f2689fff62 to your computer and use it in GitHub Desktop.

Select an option

Save itachi-re/f8f343f5bf0923b8f248d8f2689fff62 to your computer and use it in GitHub Desktop.
bugs to fix later

Waydroid Boot Loop on openSUSE Tumbleweed — Troubleshooting Log

Date: 2026-07-27 System: openSUSE Tumbleweed, kernel 7.1.3-1-default, AMD Ryzen 7 7700 (Radeon 780M iGPU), Hyprland/Wayland Waydroid version: rooted install (Magisk), home:itachi_re OBS-tracked packages Symptom: Waydroid container boots and shuts down in a repeating loop; app never reaches a usable state.


Context

This started mid-way through a pending zypper dup (224 packages, including Mesa, libdrm, kernel bump to 7.1.4). Waydroid was already broken before that upgrade was applied.


Timeline of root causes found

1. Missing ashmem_linux kernel module (ruled out as primary cause)

  • rpm -q waydroid-kmp-default showed three module builds installed (k7.1.2, k7.1.3, k7.1.4) but the running kernel was 7.1.3-1-default.
  • modprobe ashmem_linux failed: Module ashmem_linux not found in directory /usr/lib/modules/7.1.3-1-default.
  • Turned out to be expected, not a bug — this openSUSE Waydroid package explicitly sets WAYDROID_DISABLE_ASHMEM=true in its systemd override, so ashmem isn't needed. binder_linux loaded fine and was actively in use.

2. Host GPU driver lockup (host GBM passthrough) — CONFIRMED, FIXED

  • Initial waydroid_base.prop had ro.hardware.gralloc=gbm with gralloc.gbm.device=/dev/dri/renderD128 — full host GPU passthrough via Mesa GBM/EGL.
  • First crash: surfaceflinger SIGABRT with a stack trace through dri_gbm.so → libgbm_mesa.so → libEGL_mesa.so.
  • Retrying this caused a full system freeze — kernel soft lockup messages across ~10 CPUs simultaneously (surfaceflinger, RenderThread, AudioService, PowerManagerService, binder/grpc threads all stuck 100–200+ seconds). This is consistent with a GPU driver-level lock being held during a stuck submission/fence wait.
  • Recovered via Alt+SysRq REISUB (safe raw-to-shutdown key sequence for a locked box on btrfs).
  • Post-reboot: btrfs device stats showed corruption_errs: 23. A live btrfs scrub start -B / came back clean (0 errors) — the corruption was confined to systemd's own journal files getting torn mid-write by the freeze (system.journal / user-1000.journal auto-renamed and rebuilt by journald). No actual data loss. Snapshot #1890 (24 Jul) was available as a fallback but wasn't needed.
  • Fix applied: ro.hardware.gralloc=gbm → ro.hardware.gralloc=default in /var/lib/waydroid/waydroid_base.prop. This stopped the system-freezing lockups (confirmed no soft-lockup messages on subsequent boots), though Waydroid still didn't boot successfully afterward (see #4).

3. LXC container cgroup mounted read-only — CONFIRMED, FIXED

  • Every core Android service (surfaceflinger, keystore2, traced, installd, zygote, etc.) failed identically:

    libprocessgroup: Failed to make and chown /sys/fs/cgroup/uid_X: Read-only file system
    init: createProcessGroup(...) failed for service 'X': Read-only file system
    
  • Root cause: the LXC container config had lxc.mount.auto = cgroup:ro sys:ro proc — cgroup mounted read-only inside the container, so Android's init could never create per-UID cgroup subdirectories.

  • The generated session config (/var/lib/waydroid/lxc/waydroid/config) is regenerated on every session start from a template — hand-editing it directly didn't persist.

  • Source template found and patched: /usr/lib/waydroid/data/configs/config_base Changed lxc.mount.auto = cgroup:ro sys:ro proc → cgroup:rw sys:ro proc. (Backup kept at config_base.bak.)

  • Also set Delegate=yes on waydroid-container.service via systemctl edit (systemd cgroup delegation to the container's slice), though the config_base fix was the direct blocker.

  • AppArmor was checked and ruled out — zero DENIED entries in dmesg/journalctl for the relevant boot window.

  • Result: the createProcessGroup ... Read-only file system errors are completely gone after this fix. This was a real, confirmed bug.

    ⚠️ This file lives under /usr/lib/waydroid/ and will be silently reverted by any future waydroid package update. Worth re-checking after upgrades, or reporting upstream if not already a known issue — a cgroup:ro mount that breaks nearly every core Android service doesn't look intentional.

4. surfaceflinger still crashing — UNRESOLVED

  • With both fixes above in place (no GPU lockup, no cgroup errors), surfaceflinger still dies with SIGABRT in its RenderEngine thread, inside Mesa's EGL stack (libEGL_mesa.so, libGLESv2_mesa.so).
  • ro.hardware.gralloc=default changed the buffer allocator but RenderEngine still initializes an EGL context through the host Mesa driver — so Mesa hasn't actually been removed from the boot path, just touched differently.
  • Host-side coredumps (coredumpctl info surfaceflinger) only show the crash location, not the reason — stripped binaries, no useful backtrace beyond a single libc frame.
  • Attempted to pull Android's own abort reason via waydroid shell -- logcat -d, but this rooted install requires root access for the shell action and the command failed outright (ERROR: Action "shell" needs root access).
  • Not yet root-caused.

Current state

Item Status
ashmem missing Not a bug — intentionally disabled by this package
System-freezing GPU lockup Fixed — switched off host GBM passthrough
Btrfs filesystem integrity Confirmed clean — scrub, 0 errors
Cgroup read-only mount (LXC) Fixed — patched config_base, needs re-patching after Waydroid updates
systemd cgroup delegation Set to Delegate=yes (may or may not be necessary alongside the above fix)
surfaceflinger EGL/Mesa crash Still broken — root cause not yet identified

Notable side factor

This was all discovered mid-way through a pending 224-package zypper dup, which includes newer builds of exactly the libraries involved in the current crash: Mesa-dri (26.1.5-1699.2.pm.5 → pm.7), libdrm2, libdrm_amdgpu1 (2.4.134-1.2 → 2.4.134-2.1). The upgrade was never completed — there's a reasonable chance it resolves the surfaceflinger/Mesa crash outright, since it touches the exact stack that's failing. Also included: new kernel-default (7.1.4-1.1) install and a waydroid-kmp-default rebuild matched to that kernel, which would additionally resolve any ashmem/binder module-kernel version mismatch going forward.

Recommended next steps

  1. Complete the pending zypper dup now that filesystem integrity is confirmed and a snapshot (#1890, 24 Jul) exists as a rollback point. Reboot into the new 7.1.4 kernel afterward.
  2. After reboot, re-check whether /usr/lib/waydroid/data/configs/config_base still has the cgroup:rw patch (package update may have reverted it) and re-apply if needed.
  3. Re-test waydroid session start. If surfaceflinger still crashes, get root-level logcat access (may require re-rooting via Magisk against the new boot, or using waydroid shell from a session that already has root granted) to capture the actual Abort message: line — that's the missing piece for root-causing the Mesa/EGL crash.
  4. If the crash persists post-upgrade, consider testing with gralloc.gbm.device pointed at a different render node, or trying ro.hardware.gralloc=default in combination with software rendering flags (LIBGL_ALWAYS_SOFTWARE-equivalent for the container) as a diagnostic to fully rule Mesa in/out.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment