{"session": "ac564726-711b-4f3f-80d4-6576a37e879d", "env": "main-ada-funnelace", "command": "journal work end 81 --how \"after-time split measured; reported\" 2>&1|tail -1; journal helper report \"After-time: it is NOT a queue wait; it is measured from when the job starts running, and it is mostly time the job's thread is not running. (1) Code: commands/dispatch.py timed() takes started = perf_counter() INSIDE the after() closure, i.e. when the pool worker begins the job; the report is (now - started). The pool's wait is before that, in AfterAnswer.work (after_answer.py: quiet.wait_for(not held(lane), LONGEST_HOLD)) and is capped at 0.5 s anyway; queued time behind a lane's earlier job is also before job() starts. So no 'after' number quoted tonight, including the 1097 ms on the ticket merge, contains lane queueing; that part is untouched by what you feared. (2) What the job is: reply.after() = bus.background('spool') (registers and returns at once) then bus.release(queued): the handlers of the events the request raised, run inline, plus any bus_after jobs they left. (3) Measurement in-process on a copy of transportklok's record, 6 agents hooking in one environment at once, 25 s, 1344 after-jobs, load 16: wall median 73 ms, p90 110, max 5120; thread CPU median 13 ms, p90 16, max 702; lock wait (waits.total()) median 0, p90 6, max 125 ms. So of a median 73 ms, 13 is work, about 0 is lock wait, and the other 60 is the thread waiting to be scheduled (GIL contention among six threads, and the machine); the 5.1 s max has only 0.7 s of CPU and 0.1 s of lock wait. At transportklok's load of 50-70 and a 3.12 free-threaded or busy 3.14 server the same gap grows to seconds, which matches after scaling with load (median 347 ms at load 30-50, 905 ms at 50-70) and not with anything the handlers do. Answer to your question: real elapsed time of a job that runs, but almost none of it is work or lock wait; it is starvation of the thread. Do not optimise handlers against the after number; the budget row does not carry waiting or CPU for after, so a future reading cannot tell, and I suggest adding the job's thread CPU and lock wait to the after row (timing.announce already has both for the answer; it passes the after as wall only). I have not changed it. 4236 is closed (my copy 89, with Coco's outcome); I will take transportklok's 2.270.0 medians when you send them.\" 2>&1|tail -1", "parts": ["journal work end 81 --how \"after-time split measured; reported\" 2>&1|tail -1", "journal helper report \"After-time: it is NOT a queue wait; it is measured from when the job starts running, and it is mostly time the job's thread is not running. (1) Code: commands/dispatch.py timed() takes started = perf_counter() INSIDE the after() closure, i.e. when the pool worker begins the job; the report is (now - started). The pool's wait is before that, in AfterAnswer.work (after_answer.py: quiet.wait_for(not held(lane), LONGEST_HOLD)) and is capped at 0.5 s anyway; queued time behind a lane's earlier job is also before job() starts. So no 'after' number quoted tonight, including the 1097 ms on the ticket merge, contains lane queueing; that part is untouched by what you feared. (2) What the job is: reply.after() = bus.background('spool') (registers and returns at once) then bus.release(queued): the handlers of the events the request raised, run inline, plus any bus_after jobs they left. (3) Measurement in-process on a copy of transportklok's record, 6 agents hooking in one environment at once, 25 s, 1344 after-jobs, load 16: wall median 73 ms, p90 110, max 5120; thread CPU median 13 ms, p90 16, max 702; lock wait (waits.total()) median 0, p90 6, max 125 ms. So of a median 73 ms, 13 is work, about 0 is lock wait, and the other 60 is the thread waiting to be scheduled (GIL contention among six threads, and the machine); the 5.1 s max has only 0.7 s of CPU and 0.1 s of lock wait. At transportklok's load of 50-70 and a 3.12 free-threaded or busy 3.14 server the same gap grows to seconds, which matches after scaling with load (median 347 ms at load 30-50, 905 ms at 50-70) and not with anything the handlers do. Answer to your question: real elapsed time of a job that runs, but almost none of it is work or lock wait; it is starvation of the thread. Do not optimise handlers against the after number; the budget row does not carry waiting or CPU for after, so a future reading cannot tell, and I suggest adding the job's thread CPU and lock wait to the after row (timing.announce already has both for the answer; it passes the after as wall only). I have not changed it. 4236 is closed (my copy 89, with Coco's outcome); I will take transportklok's 2.270.0 medians when you send them.\" 2>&1|tail -1"]}