Date: 2026-07-27
System: openSUSE Tumbleweed, kernel 7.1.3-1-default, AMD Ryzen 7 7700 (Radeon 780M iGPU), Hyprland/Wayland
Waydroid version: rooted install (Magisk), home:itachi_re OBS-tracked packages
Symptom: Waydroid container boots and shuts down in a repeating loop; app never reaches a usable state.
This started mid-way through a pending zypper dup (224 packages, including Mesa, libdrm, kernel bump to 7.1.4). Waydroid was already broken before that upgrade was applied.
rpm -q waydroid-kmp-defaultshowed three module builds installed (k7.1.2,k7.1.3,k7.1.4) but the running kernel was7.1.3-1-default.modprobe ashmem_linuxfailed:Module ashmem_linux not found in directory /usr/lib/modules/7.1.3-1-default.- Turned out to be expected, not a bug — this openSUSE Waydroid package explicitly sets
WAYDROID_DISABLE_ASHMEM=truein its systemd override, so ashmem isn't needed.binder_linuxloaded fine and was actively in use.
- Initial
waydroid_base.prophadro.hardware.gralloc=gbmwithgralloc.gbm.device=/dev/dri/renderD128— full host GPU passthrough via Mesa GBM/EGL. - First crash:
surfaceflingerSIGABRT with a stack trace throughdri_gbm.so→libgbm_mesa.so→libEGL_mesa.so. - Retrying this caused a full system freeze — kernel
soft lockupmessages across ~10 CPUs simultaneously (surfaceflinger, RenderThread, AudioService, PowerManagerService, binder/grpc threads all stuck 100–200+ seconds). This is consistent with a GPU driver-level lock being held during a stuck submission/fence wait. - Recovered via
Alt+SysRqREISUB (safe raw-to-shutdown key sequence for a locked box on btrfs). - Post-reboot:
btrfs device statsshowedcorruption_errs: 23. A livebtrfs scrub start -B /came back clean (0 errors) — the corruption was confined to systemd's own journal files getting torn mid-write by the freeze (system.journal/user-1000.journalauto-renamed and rebuilt by journald). No actual data loss. Snapshot#1890(24 Jul) was available as a fallback but wasn't needed. - Fix applied:
ro.hardware.gralloc=gbm→ro.hardware.gralloc=defaultin/var/lib/waydroid/waydroid_base.prop. This stopped the system-freezing lockups (confirmed no soft-lockup messages on subsequent boots), though Waydroid still didn't boot successfully afterward (see #4).
-
Every core Android service (
surfaceflinger,keystore2,traced,installd,zygote, etc.) failed identically:libprocessgroup: Failed to make and chown /sys/fs/cgroup/uid_X: Read-only file system init: createProcessGroup(...) failed for service 'X': Read-only file system -
Root cause: the LXC container config had
lxc.mount.auto = cgroup:ro sys:ro proc— cgroup mounted read-only inside the container, so Android'sinitcould never create per-UID cgroup subdirectories. -
The generated session config (
/var/lib/waydroid/lxc/waydroid/config) is regenerated on every session start from a template — hand-editing it directly didn't persist. -
Source template found and patched:
/usr/lib/waydroid/data/configs/config_baseChangedlxc.mount.auto = cgroup:ro sys:ro proc→cgroup:rw sys:ro proc. (Backup kept atconfig_base.bak.) -
Also set
Delegate=yesonwaydroid-container.serviceviasystemctl edit(systemd cgroup delegation to the container's slice), though theconfig_basefix was the direct blocker. -
AppArmor was checked and ruled out — zero
DENIEDentries indmesg/journalctlfor the relevant boot window. -
Result: the
createProcessGroup ... Read-only file systemerrors are completely gone after this fix. This was a real, confirmed bug.⚠️ This file lives under/usr/lib/waydroid/and will be silently reverted by any futurewaydroidpackage update. Worth re-checking after upgrades, or reporting upstream if not already a known issue — acgroup:romount that breaks nearly every core Android service doesn't look intentional.
- With both fixes above in place (no GPU lockup, no cgroup errors),
surfaceflingerstill dies withSIGABRTin itsRenderEnginethread, inside Mesa's EGL stack (libEGL_mesa.so,libGLESv2_mesa.so). ro.hardware.gralloc=defaultchanged the buffer allocator but RenderEngine still initializes an EGL context through the host Mesa driver — so Mesa hasn't actually been removed from the boot path, just touched differently.- Host-side coredumps (
coredumpctl info surfaceflinger) only show the crash location, not the reason — stripped binaries, no useful backtrace beyond a single libc frame. - Attempted to pull Android's own abort reason via
waydroid shell -- logcat -d, but this rooted install requires root access for theshellaction and the command failed outright (ERROR: Action "shell" needs root access). - Not yet root-caused.
| Item | Status |
|---|---|
| ashmem missing | Not a bug — intentionally disabled by this package |
| System-freezing GPU lockup | Fixed — switched off host GBM passthrough |
| Btrfs filesystem integrity | Confirmed clean — scrub, 0 errors |
| Cgroup read-only mount (LXC) | Fixed — patched config_base, needs re-patching after Waydroid updates |
| systemd cgroup delegation | Set to Delegate=yes (may or may not be necessary alongside the above fix) |
| surfaceflinger EGL/Mesa crash | Still broken — root cause not yet identified |
This was all discovered mid-way through a pending 224-package zypper dup, which includes newer builds of exactly the libraries involved in the current crash: Mesa-dri (26.1.5-1699.2.pm.5 → pm.7), libdrm2, libdrm_amdgpu1 (2.4.134-1.2 → 2.4.134-2.1). The upgrade was never completed — there's a reasonable chance it resolves the surfaceflinger/Mesa crash outright, since it touches the exact stack that's failing. Also included: new kernel-default (7.1.4-1.1) install and a waydroid-kmp-default rebuild matched to that kernel, which would additionally resolve any ashmem/binder module-kernel version mismatch going forward.
- Complete the pending
zypper dupnow that filesystem integrity is confirmed and a snapshot (#1890, 24 Jul) exists as a rollback point. Reboot into the new7.1.4kernel afterward. - After reboot, re-check whether
/usr/lib/waydroid/data/configs/config_basestill has thecgroup:rwpatch (package update may have reverted it) and re-apply if needed. - Re-test
waydroid session start. If surfaceflinger still crashes, get root-levellogcataccess (may require re-rooting via Magisk against the new boot, or usingwaydroid shellfrom a session that already has root granted) to capture the actualAbort message:line — that's the missing piece for root-causing the Mesa/EGL crash. - If the crash persists post-upgrade, consider testing with
gralloc.gbm.devicepointed at a different render node, or tryingro.hardware.gralloc=defaultin combination with software rendering flags (LIBGL_ALWAYS_SOFTWARE-equivalent for the container) as a diagnostic to fully rule Mesa in/out.