Three fixes in input_ibus.c, found by capturing the IBus session bus while a
native Mac RDP client and a Guacamole (browser) client typed Chinese.
1. Keep the IBus thread alive across client disconnects.
xrdp_input_unicode_destroy() used to stop the IBus thread and its private
GMainContext on every disconnect, and unicode_init() started a new pair on
every reconnect. ibus_bus_new() returns a process-wide singleton whose
GDBusConnection stays bound to the context that was thread-default when it
was first created, so a reconnect got the old connection back, bound to a
context nobody iterates. ibus-daemon's CreateEngine call on our factory was
never dispatched and every ibus_bus_set_global_engine() timed out ("failed
to switch global engine to XrdpIme", 15 s each), losing all Unicode input
until the session was fully logged out. (Same connection name on both
connects in the bus trace.) destroy() now only restores the user's engine
and drains the queue on the IBus thread (xrdp_input_reset_cb) and waits,
bounded, for that; init() takes its existing "already ready" path.
2. Do not re-request XrdpIme while an instance is being created.
Two characters arriving ~30 us apart each called set_global_engine(): the
second found the name already XrdpIme but no instance yet, fell through,
and made ibus destroy the instance under construction and build another
("IM enabled / IM disabled / IM enabled"), dropping what was queued for the
first. A request made within XRDP_INPUT_ENGINE_CREATE_WAIT_US (2 s) is now
treated as in flight and left alone; an older one is still treated as a
stale name and re-set.
3. Measure the commit debounce from when the text was queued.
"Quiet" was measured from the last raw key only. A user who pauses between
typing the pinyin and confirming it (270 ms in the capture) makes the raw
keys already quiet when the commit arrives, so it was committed at once and
the client's cleanup Backspaces, one per leaked key and arriving 1-2 ms
later, deleted the committed text and kept going into what was there
before. Quiet is now measured from the later of the last raw key and the
queue time, so every commit waits at least XRDP_INPUT_COMMIT_QUIET_US and
any Backspace inside the window extends it (still capped by MAX_WAIT).
Tested against a real xrdp session in the agent-cn image, through guacd with a
script that mimics the native client (raw keys, pause, Unicode commit,
Backspaces after a variable delay): with the old code any delay >= 8 ms lost
the commit ("我是" became "wo"); now delays up to 90 ms pass, 110 ms and up
still fail. A fresh session and a reconnect both commit correctly, and the
"failed to switch global engine" timeouts are gone. Confirmed working with the
native client by hand. Costs every Unicode commit 100 ms of latency.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Microsoft Remote Desktop for Mac was observed to garble Chinese text
typed through its local Pinyin IME: the composed characters would
land with pieces missing or extra characters deleted from the
document. Root cause, confirmed with debug logging plus dbus-monitor
on the live ibus session bus: the client forwards the raw
pre-composition keystrokes over the normal RDP keyboard channel while
composing (these reach XrdpIme's process-key-event handler and, since
it never consumes anything, get typed into the document as literal
ASCII), then once composition finishes sends a compensating burst of
Backspace keystrokes sized to exactly undo them, interleaved with the
final composed text arriving over the separate TS_UNICODE_KEYBOARD_EVENT
side channel. The two channels have no ordering guarantee, so the
side-channel commit could land in the middle of the client's own
backspace cleanup and get partially deleted.
Measured the backspace burst against the leaked raw-character count
across several phrases (23-for-23, 29-for-29) - it's always exactly
right, so this is a pure interleaving race, not a client miscount.
That makes it fixable server-side: xrdp_input_queue_source_prepare/
_check now withhold a queued commit until raw key traffic on the
engine has been quiet for XRDP_INPUT_COMMIT_QUIET_US (100ms), so the
client's own cleanup has time to finish first, capped by
XRDP_INPUT_COMMIT_MAX_WAIT_US (500ms) so back-to-back composition
can't delay a commit indefinitely. This is an empirical mitigation
tuned to observed client behaviour, not a protocol guarantee.
Also fixes two latent data races found while reviewing this code
against the cross-thread invariants already documented at the top of
the file: g_engine and last_input_name are written by the IBus thread
(engine enable/disable, an unexpected daemon disconnect) but were read
and, for last_input_name, also written directly by xrdp_input_enable()
on chansrv's own thread with no synchronization - unlike `bus`, which
already gets this treatment. Both now follow the same
publish/clear-under-state_mutex pattern as `bus`.
Added IBUS-KEY/IBUS-COMMIT debug logging used to diagnose this.
Verified live against a real Microsoft Remote Desktop for Mac client
connected to an xrdp/ibus session: multiple Chinese phrases of varying
length composed and committed cleanly with the debounce in place,
after reliably reproducing the interleaved garbling without it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The "don't reclaim a deliberately-chosen engine" guard from the last
commit had a logic bug that broke ordinary Model A composition
completely, not just the edge case it was meant to fix.
ibus_bus_set_global_engine() sets the global engine *name*
synchronously, but the actual engine instance (g_engine) is created
asynchronously afterward. During that brief window, the global engine
name is already "XrdpIme" but g_engine is still NULL - the existing
fast path (name == "XrdpIme" && g_engine != NULL) correctly falls
through in that case, same as always. But the new guard only checked
whether the current name differed from last_input_name, and "XrdpIme"
is *always* different from last_input_name by construction (that's
the name of whatever we switched away from). So every commit that
landed in that async window got misclassified as "a different engine
took over" and silently dropped. Composing even one phrase sends
several characters in quick succession, so this was most of them -
not just the intended edge case.
Fixed by checking name == "XrdpIme" first, independent of g_engine's
readiness, before ever reaching the "is this some other engine"
guard - that window means "still connecting", not "something else
took over".
Verified against the exact failure condition: sent a 9-character
phrase rapid-fire immediately after xrdp_input_unicode_init() (no
settle time), landing correctly as a single unbroken phrase with zero
drops. Also re-ran all four trigger-happy scenarios from the previous
commit, plus a new check sending 5 characters rapid-fire immediately
after a from-original reassertion - all pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
xrdp_input_enable() unconditionally reasserted XrdpIme as the global
engine any time it wasn't already active, on the theory that ibus can
silently revert to the user's original engine on its own. That's true,
but the fix was too broad: it also fired whenever the user had
manually switched the session's engine to something else entirely
(e.g. libpinyin, to compose Chinese directly in the remote desktop
rather than via the client's local IME), yanking control away
mid-use. Confirmed happening live: a composition left pending in a
client-side IME and then abandoned got auto-flushed by the OS well
after the fact, silently reclaiming XrdpIme and interrupting an
unrelated libpinyin session already in progress.
Now only reclaims XrdpIme when the current engine is unset or matches
the original one remembered before ever switching away from it
(last_input_name) - the actual "ibus reverted on its own" case this
existed for. Any other named engine is treated as a deliberate choice
and left alone; the stale/deferred commit is dropped instead of
hijacking whatever the user is currently doing.
Also fixes a hang found while testing this: if xrdp_input_unicode_init()
times out waiting for ibus_get_address() (daemon unreachable), it
returned early without resetting ibus_thread_exited back to TRUE, even
though it had just set it FALSE in anticipation of a thread that was
never actually created. No thread means no broadcast ever comes, so
any later xrdp_input_unicode_destroy() or re-init() call blocks on
that condvar forever. Pre-existing in the private-loop rewrite, not
introduced by this change - just surfaced by testing this one harder.
Verified: 20 back-to-back init() calls on a healthy connection still
take the fast path with no false failures; manually switching to a
different engine and then sending a stale commit leaves that engine
untouched and fails the commit; reverting to the original engine still
correctly reasserts XrdpIme; 5 kill/restart cycles still show no
thread leak; an 18-char rapid-fire send against a real focused GTK
entry still lands every character in order.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three correctness gaps found in code review of the private-loop
rewrite, all in how chansrv's thread and the IBus thread interact
around `bus`:
- xrdp_input_enable() read the shared `bus` pointer with no
synchronization at all, despite the IBus thread being free to null
it out (disconnect, teardown) at any time - a real use-after-free
risk, not just a formality, since this function actively calls
methods on it rather than just checking it's non-NULL. Fixed by
taking a ref under state_mutex before touching it, and using that
local reference for the rest of the call; bus's creation and
teardown in xrdp_input_main_loop() now publish/clear the shared
pointer under the same lock so the ref-under-lock pattern actually
synchronizes against something.
- xrdp_input_unicode_init()'s fast path trusted ibus_ready without
re-checking the connection. ibus_ready and connection liveness
normally change together, but there's a window between the
underlying socket actually dying and the "disconnected" signal being
dispatched on the IBus thread. A client that reconnects into that
window would get a false "unicode input is supported"
advertisement it could never actually use. Now double-checks
ibus_bus_is_connected() before taking the fast path.
- That staleness check can't just flip ibus_ready and fall through to
spawning a new thread - the old one may still be alive and about to
touch bus/g_engine itself, racing the new thread's own setup. Now
asks the stale thread to shut down (g_main_loop_quit, safe to call
cross-thread) and waits on the exit condvar before proceeding, so
there's never more than one IBus thread alive at a time.
Verified: 20 back-to-back xrdp_input_unicode_init() calls on a healthy
connection all take the fast path correctly (no spurious thread
restarts); 5 kill/restart cycles still show no thread leak; an 18-char
rapid-fire send against a real focused GTK entry still lands every
character in order with zero drops after these changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
XrdpIme never actually became the active engine after the private
main loop rewrite - ibus_bus_set_global_engine() returned FALSE every
time. Confirmed the actual D-Bus error by bypassing the boolean-only
wrapper and calling SetGlobalEngine directly:
Set global engine failed: Timeout was reached
Root cause: xrdp_input_enable() was being marshaled onto the IBus
thread via g_main_context_invoke(), on the (wrong) theory that it
"must run on the IBus thread" like commit_text does. But
set_global_engine() blocks synchronously waiting for ibus-daemon's
reply, and ibus-daemon can't reply until it finishes instantiating the
engine - which means calling back into our own IBusFactory's
"create-engine" over the same connection. Running xrdp_input_enable()
on the IBus thread left that thread stuck blocked inside its own
synchronous call, unable to service the nested callback ibus-daemon
needed answered before it could reply - a self-deadlock, resolved only
by ibus-daemon's own timeout.
Calling xrdp_input_enable() directly from chansrv's thread instead
(matching the original pre-rewrite code, which did this without
apparent reason but happened to sidestep the deadlock) leaves the IBus
thread free to answer the nested call while chansrv's thread blocks.
GDBusConnection sync calls are documented thread-safe to issue from
any thread, so this doesn't need marshaling despite `bus` otherwise
being owned by the IBus thread.
Verified end-to-end against a real ibus-daemon and a focused GTK entry
widget (not just the D-Bus trace): ibus engine now correctly reports
"XrdpIme" after a send, and 27 characters fired with zero delay
between them - the exact rapid-fire pattern that used to drop most
commits under the old design - landed in order with zero drops and
zero duplicates.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replace the shared-global-context design with one where chansrv's
thread never touches libibus directly. All IBus objects (bus, engine,
factory, component) are now owned exclusively by a private
GMainContext/GMainLoop created inside the IBus thread; chansrv hands
off unicode codepoints through a thread-safe GAsyncQueue and wakes the
loop, instead of calling ibus_engine_commit_text() from the wrong
thread or relying on ibus_main()/ibus_quit(), which turned out to
operate on a single loop shared by the whole process rather than one
per connection.
Fixes along the way, found by testing against a live ibus-daemon and
a real focused GTK app rather than just reading the diff:
- xrdp_engine_enable()/disable() now take a real reference on g_engine
(g_object_ref/unref) instead of caching a bare pointer. The old code
unreffed g_engine on teardown without ever having reffed it - IBusEngine
derives from GInitiallyUnowned, so that was releasing a reference it
never owned.
- A custom GSource drains the unicode queue, gated on the engine
actually existing (engine creation is async once XrdpIme is
selected), so a character queued before the engine is ready isn't
lost - it commits as soon as the "enable" signal fires, instead of
racing a one-shot commit against that async setup.
- ibus_engine_commit_text() releases its IBusText argument itself
(it's floating, documented behavior for that specific call) - an
earlier draft of this rewrite added an extra g_object_unref() after
it, which would have been a double-release.
- bus (and anything else that opens a GDBusConnection) must be created
after the private GMainContext is pushed as thread-default, not
before - otherwise its async I/O silently binds to the global
default context, which nothing here iterates, and signals like
"disconnected" simply never fire.
- xrdp_input_unicode_init() now blocks on a condvar until the IBus
thread actually signals ready (or exits), instead of returning
immediately and hoping the engine shows up in time.
- unicode_init()/destroy() no longer touch bus/g_engine from chansrv's
own thread while the IBus thread may still be running against them;
destroy() schedules real teardown on the IBus thread and joins it via
a condvar before returning.
Known limitation, confirmed empirically rather than assumed: this does
NOT make reconnect-after-daemon-restart (#3230) actually work.
ibus_bus_new() returns a process-wide singleton in this libibus
version - two calls in the same process with no disconnect involved
return the identical pointer - so once its connection dies, every
later call just hands back the same dead object; ibus_bus_is_connected()
on it stays permanently false, confirmed not to be a transient state
via repeated retries with delay. The "disconnected" handler still
tears the thread down cleanly so a later xrdp_input_unicode_init()
fails fast instead of hanging or operating on stale state, but a real
fix would mean bypassing IBusBus for a raw GDBusConnection to
ibus-daemon, which isn't justified given how rarely the daemon
actually restarts mid-session.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A few follow-on fixes found while testing the ibus reconnect changes
in a live session:
- xrdp_input_unicode_init() failed outright if ibus had no default
global engine yet at connect time, which is a normal state on a
fresh session (nothing has chosen one yet), not an error. Since
nothing else used the return value, this permanently killed unicode
input for the whole session over a spurious check.
- ibus can disable our engine instance (e.g. on a focus change) while
leaving the global engine name as "XrdpIme", since those are tracked
separately. xrdp_input_enable()'s fast path only checked the name,
so it could skip re-asserting and leave xrdp_input_send_unicode()
committing text through a disabled g_engine. Clear g_engine on
disable and require it to be set for the fast path to apply.
- ibus_engine_commit_text() was being called from chansrv's own
thread, not the thread pumping the glib main loop that owns the
engine's D-Bus connection (xrdp_input_main_loop). Marshal the actual
commit through g_main_context_invoke() onto the correct thread.
None of these are the full fix for intermittent dropped commits during
real pinyin input testing - that's still open, with the current lead
being ibus Reset calls interleaved with commit bursts, likely from the
FocusOut/FocusIn churn caused by passing every raw keystroke through
engine_process_key_event_cb. To be continued.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Fixes#3230. The static `bus` global was only ever checked for
non-NULL, not for whether the underlying connection was still alive.
If ibus disconnected (daemon restart, stale socket) or the initial
connect attempt failed, `bus` was left set to a dead/freed connection,
so every later call to xrdp_input_unicode_init() took the "already
initialized" fast path and operated on it.
- Null out bus/g_engine in the "disconnected" signal handler instead
of leaving them dangling after g_object_unref().
- Check ibus_bus_is_connected() before trusting a cached bus, and
tear down + reconnect if it's stale.
- Unref and clear bus on a failed connect attempt instead of leaving
it set.
- Guard the unrefs in xrdp_input_unicode_destroy() now that bus/
g_engine can legitimately already be NULL.
The initial implementation of Uinicode input via IBus used a startup delay
of 3 seconds to wait for the daemon to be ready before connecting to it.
This commit introduces a poll-wait loop which can remove the delay
entirely if the daemon is up when chansrv starts the interface.