fix(data): cleanup uses failed candidate status instead of 504

The stale-pending cleanup task previously hardcoded status_code=504 and a
generic timeout message for every usage row it finalized. When a request
had already been observed as failing — e.g. upstream Connection reset by
peer, watchdog 504, or an authenticated 4xx — the cleanup overwrote that
context with a misleading "服务器超时" outcome and 504 status, hiding the
real cause from the dashboards and customer.

Pull the most recent failed/cancelled candidate per stale request_id and,
if present, finalize the usage row with the candidate's status_code
(defaulting to 502 when none was recorded) and error_message. Requests
that have no terminal candidate (truly stuck pending/streaming) keep the
existing 504 + timeout-message behavior, since they really are timeouts
from the cleanup's perspective. Applied to all three SQL backends with
parameterized UPDATE statements.

The Postgres failed-candidate lookup orders by
COALESCE(finished_at, started_at, created_at) DESC, matching the MySQL
and SQLite ORDER BY clauses so the three backends pick the same
"most recent terminal candidate" under every NULL combination of timing
columns.

Co-Authored-By: Claude Opus 4.7 <[email protected]>
This commit is contained in:
stabey
2026-05-21 16:56:56 +08:00
co-authored by Claude Opus 4.7
parent 40434005c0
commit cadc45c5b8
4 changed files with 315 additions and 14 deletions
@@ -825,6 +825,14 @@ fn usage_sql_clears_stale_failure_fields_for_non_failed_status_updates() {
));
}
#[test]
fn stale_cleanup_failed_candidate_sql_orders_by_effective_timestamp() {
let sql = super::SELECT_LATEST_FAILED_CANDIDATE_FOR_STALE_REQUESTS_SQL;
assert!(sql.contains("COALESCE(finished_at, started_at, created_at) DESC"));
assert!(!sql.contains("finished_at DESC NULLS LAST"));
assert!(!sql.contains("started_at DESC NULLS LAST"));
}
#[test]
fn usage_sql_does_not_allow_streaming_to_regress_back_to_pending() {
assert!(super::UPSERT_SQL.contains(