몇 주 동안 Elevare Digital의 크론 잡(cron job)은 예정된 시간에 깨어나 큐를 확인하고 성공 로그를 남겼습니다. 하지만 승인된 초안은 정확히 0건이었습니다. 19개의 콘텐츠가 대기 상태로 남아 있었습니다. 팀은 나중에야, 이 침묵의 공백이 단순한 이상 현상을 넘어 작은 백로그로 커진 후에야 이 사실을 알게 되었습니다. 아무것도 충돌하지 않았습니다. 페이징 알람도 울리지 않았습니다. 시스템은 기술적으로는 정상이었으나, 기능적으로는 마비된 상태였습니다.
이것이 자율형 파이프라인이 가진 조용한 공포입니다. 루프에서 인간을 제거하면, 아무 일도 일어나지 않고 있다는 사실을 알아차릴 사람마저 제거하게 됩니다.
스스로 돌아가는 파이프라인
Elevare Digital은 완전히 자동화된 콘텐츠 워크플로우를 운영합니다. 소프트웨어 에이전트가 초안을 생성합니다. 예약된 승인 크론(approver cron)이 게이트키퍼 역할을 하며, 해당 초안을 검토하고 승인된 항목을 즉시 발행 단계로 넘깁니다. 사람이 대시보드를 열어 각 배치를 승인하는 일은 없습니다. 기계가 번거로운 작업을 처리하는 동안 팀은 다른 문제에 집중할 수 있게 하는 것이 이 시스템의 핵심입니다.
이 모델에서는 신뢰가 주요 인터페이스가 됩니다. 스케줄러가 실행될 것을 신뢰합니다. 작업이 돌아갈 것을 신뢰합니다. 종료 코드(exit code)를 신뢰합니다. 로그에 200 OK 응답이 꾸준히 찍히면 작업이 진행되고 있다고 가정합니다. 몇 주 동안 그 심박수는 완벽했습니다. 크론은 매번 제시간에 실행되었습니다. 다만 실제 작업은 전혀 하지 않았을 뿐입니다.
19개의 초안과 무음의 경보
발견은 우연히 이루어졌습니다. 누군가 발행 큐가 조용해졌다는 것을 눈치챘거나, 혹은 다운스트림 지표를 확인하다가 그래프가 평탄해진 것을 발견했습니다. 그들이 찾아낸 것은 전혀 손대지 않은 채 쌓여 있는 19개의 초안이었습니다. 승인 프로세스는 매일 성실하게 실행되며 성공 로그를 남겼지만, 단 하나의 초안도 처리하지 못했습니다.
수동 워크플로우였다면, 검토자가 첫날부터 비어 있는 편지함이나 쌓여 있는 대기 항목을 발견했을 것입니다. 하지만 자동화된 버전에서는 활동의 부재가 작업의 부재와 똑같이 보였습니다. 크론에게는 실망을 줄 관리자가 없었습니다. 그저 출근 도장만 찍고 조기 퇴근을 반복했을 뿐입니다.
두 개의 버그, 하나의 빈 결과
이 실패에는 두 가지 원인이 있었습니다. 둘 다 구문 오류(syntax error)나 타임아웃, 혹은 의존성 장애가 아니었습니다. 둘 다 쿼리 엔진의 관점에서 19개의 유효한 행을 아무것도 없는 상태로 만들어버린 의미론적(semantic) 실수였습니다.
첫째, 타입 불일치(type mismatch)입니다. 초안을 생성하는 에이전트는 레코드를 article로 태그하여 작성했습니다. 하지만 승인 크론은 구체적으로 thread 타입만을 조회하도록 설정되어 있었습니다. 이는 생산자와 소비자가 서로 다른 경로로 발전할 때 발생하는 전형적인 드리프트(drift) 현상입니다. 한 팀(또는 에이전트)은 출력이 article이라고 결정했고, 다른 팀은 소비자가 thread를 가져올 것이라고 가정하고 코드를 작성했습니다. 이들은 아마도 느슨한 문자열 태그, JSON 필드 또는 제약이 없는 varchar 값이었기에 타입 시스템에서 컴파일 타임 에러를 던지지 않았습니다. 데이터베이스는 단순히 일치하는 항목을 찾지 못해 빈 세트(empty set)를 반환했습니다. 엔진 입장에서 이것은 에러 상황이 아닙니다. 잘못된 질문에 대한 올바른 답변일 뿐입니다.
둘째, 승인자 쿼리의 인너 조인(inner join)이 행들을 조용히 삼켜버렸습니다. 만약 쿼리가 초안 테이블을 다른 테이블(메타데이터, 상태 플래그 또는 라우팅 규칙 조회를 위한 테이블 등)과 조인하는데, 조인 조건이 일치하지 않는다면 인너 조인은 설계된 대로 정확하게 동작합니다. 일치하지 않는 행을 제외해 버리는 것입니다. 결과 세트에 고아 행(orphan rows)이 나타나지도 않았고, null 값이 문제를 알리지도 않았습니다. 19개의 초안은 체를 통과하는 물처럼 쿼리를 빠져나갔고, 애플리케이션 레이어는 깨끗하게 비어 있는 리스트를 전달받았습니다.
쿼리가 아무런 행도 반환하지 않았기 때문에 함수는 깔끔하게 종료되었습니다. 예외(exception)가 발생하지 않았습니다. HTTP 응답은 200 OK였습니다. 크론은 성공을 기록하고 다시 잠들었습니다.
'0건 처리'의 함정
여기에 문제의 핵심이 있습니다. 큐 기반 시스템에서 컨슈머는 처리할 행이 0개인 상황을 자주 마주합니다. 큐가 비어 있고, 작업자가 빠르게 작업을 마칩니다. 로그에는 processed: 0이라고 찍히며, 팀은 이를 "수요를 잘 따라가고 있다"는 좋은 소식으로 해석합니다. 이는 정상적인 상태입니다.
하지만 processed: 0은 완전히 다른 두 가지 현실을 내포합니다.
- 정상 상태: 대기 중인 작업이 0개라서 처리량이 0개임. 큐가 비어 있음. 시스템이 설계된 대로 유휴(idle) 상태임.
- 고장 난 상태: 컨슈머가 작업을 볼 수 없어서 처리량이 0개임. 큐에는 19개의 행이 있음. 시스템이 유휴 상태가 아니라 눈이 먼 상태임.
큐의 깊이(queue depth)에 대한 독립적인 확인 없이는, 이 두 상태는 동일한 텔레메트리(telemetry)를 내보냅니다. 대시보드에서는 똑같이 보이고, 로그 애그리게이터에서도 똑같은 냄새가 나며, PagerDuty에서도 똑같은 침묵을 유발합니다. 당신은 작업자가 비명을 지를 때를 감지하는 모니터링 전략을 구축했을 뿐, 실제 작업 더미 옆에서 속삭이며 지나갈 때는 감지하지 못하는 시스템을 만든 것입니다.
간극 메우기
Elevare Digital fixed the problem by changing what they monitor. They stopped relying solely on error rates and success statuses. Instead, they started alerting on the gap between available work and completed work.
After every batch, they now run a simple invariant check:
- If processed is 0 and pending rows are greater than 0, trigger a high severity alert.
This rule is deliberately agnostic about cause. It does not care if the miss was a bad filter, a broken join, or a mistyped enum string. It cares only that work exists and no work got done. This shifts monitoring from “Did the process complain?” to “Did the work move?”
To support this, they treat queue depth as a first-class metric, tracked over time, not just as a spot-check. If the producer keeps adding rows while the consumer continuously reports success, the depth trend turns into a smoking gun. A static snapshot might lie, but a creeping backlog never does.
Lessons for Autonomous Systems
The Elevare incident contains a handful of practical rules for anyone running hands-off pipelines.
Log scanned rows separately from processed rows. The consumer might execute a query that touches forty rows, filters them all out through bad criteria, and reports processed: 0. If you only log the final count, you miss the ghost interaction. A scanned-rows metric reveals that the worker showed up, looked at the work, and walked away confused. That gap between scanned and processed is often your earliest signal.
Track queue depth as a time-series. A queue that is temporarily empty is fine. A queue that grows monotonically while workers stay green is not. Plot depth against consumer throughput. When the two diverge, investigate immediately, even if every health check is passing.
Test consumers against real producer output, not just mocks. Unit tests with mocked data carry the assumptions of the tester. If the mock factory produces thread types and the consumer expects thread types, your tests pass while production fails. Run integration tests that pull actual records from the producer’s output. Make sure the consumer can truly see what the producer writes.
Treat data types and enum values as contracts. Loose string tags in JSON blobs are convenient until they become invisible failure points. Define schemas explicitly. Share constants. Validate payloads at the seam between producer and consumer. If the contract breaks, the system should fail loudly at the boundary, not silently inside a WHERE clause.
The Real Takeaway
Autonomous systems do not fail like humans. They do not call in sick, throw exceptions every time, or leave obvious crash dumps. They return 200 OK and let the inventory rot. If your alerts only listen for screams, you will miss the most expensive failures—the ones where everything looks fine and nothing gets done.
Design your observability to watch the gap. Measure the work that enters against the work that exits. When the two no longer match, assume the machine is lying to you. Because sometimes, a perfect success log is the only symptom of a system that has gone completely blind.
