Android’s build tools just got a speed-boost for Kotlin Coroutines. With the Android Gradle Plugin (AGP) 9.2.0 and the R8 shrinker bundled inside, the bytecode that powers coroutine state handling is rewritten to deliver roughly twice the performance in the most common coroutine operations. Developers who ship Android apps that rely on coroutines can see faster cold-start times, steadier frame rates and a modest improvement in battery life without changing a single line of Kotlin code.

เครื่องมือการ Build ของ Android เพิ่งได้รับการเพิ่มความเร็วสำหรับ Kotlin Coroutines ด้วย Android Gradle Plugin (AGP) 9.2.0 และ R8 shrinker ที่รวมมาให้ในตัว โดย bytecode ที่ใช้จัดการสถานะของ coroutine จะถูกเขียนขึ้นใหม่เพื่อให้ได้ประสิทธิภาพเพิ่มขึ้นประมาณสองเท่าในการทำงานของ coroutine ที่พบบ่อยที่สุด นักพัฒนาที่ส่งแอป Android ที่พึ่งพา coroutines จะเห็นเวลาในการ cold-start ที่เร็วขึ้น อัตราเฟรม (frame rates) ที่นิ่งขึ้น และอายุการใช้งานแบตเตอรี่ที่ปรับปรุงขึ้นเล็กน้อย โดยไม่ต้องเปลี่ยนโค้ด Kotlin แม้แต่บรรทัดเดียว

Why coroutine performance mattered

ทำไมประสิทธิภาพของ coroutine ถึงสำคัญ

Kotlin Coroutines use AtomicFieldUpdater objects to modify internal fields safely from multiple threads. The approach avoids allocating a new atomic object for every update, but it introduces three hidden costs on Android:

Kotlin Coroutines ใช้ AtomicFieldUpdater เพื่อแก้ไขฟิลด์ภายในอย่างปลอดภัยจากหลายเธรด วิธีนี้ช่วยหลีกเลี่ยงการสร้างออบเจกต์ atomic ใหม่ในทุกการอัปเดต แต่ก็นำมาซึ่งต้นทุนแฝง 3 ประการบน Android:

  • Class-loading delay – the updater is created reflectively, so the VM must look up the target field at runtime.
  • ความล่าช้าในการโหลดคลาส (Class-loading delay) – ตัว updater ถูกสร้างขึ้นผ่านการสะท้อน (reflection) ดังนั้น VM จึงต้องค้นหาฟิลด์เป้าหมายในขณะรันไทม์
  • CPU waste – each access checks the updater’s state before reaching the actual field.
  • การสิ้นเปลือง CPU – ทุกการเข้าถึงจะต้องตรวจสอบสถานะของ updater ก่อนที่จะเข้าถึงฟิลด์จริง
  • Inlining limits – the compiler cannot inline the updater calls, keeping the generated bytecode larger and slower.
  • ข้อจำกัดในการทำ Inlining – คอมไพเลอร์ไม่สามารถทำ inline การเรียกใช้งาน updater ได้ ทำให้ bytecode ที่สร้างขึ้นมีขนาดใหญ่ขึ้นและช้าลง

In practice these costs surface during hot paths such as channel send/receive, mutex lock/unlock, StateFlow updates and coroutine dispatch. The result is a noticeable amount of CPU time spent on bookkeeping rather than on the app’s own work.

ในทางปฏิบัติ ต้นทุนเหล่านี้จะปรากฏขึ้นในช่วง hot paths เช่น การส่ง/รับข้อมูลผ่าน channel, การ lock/unlock ของ mutex, การอัปเดต StateFlow และการ dispatch coroutine ผลลัพธ์ที่ได้คือเวลาของ CPU จำนวนมากถูกใช้ไปกับการจัดการข้อมูลเบื้องหลัง (bookkeeping) แทนที่จะเป็นงานหลักของแอป

What changed in AGP 9.2.0 and R8

มีอะไรเปลี่ยนไปใน AGP 9.2.0 และ R8

R8, the code-shrinker that ships with AGP 9.2.0, now scans the compiled bytecode for the standard AtomicFieldUpdater pattern. When it finds one, it performs four transformations:

R8 ซึ่งเป็นตัวย่อโค้ด (code-shrinker) ที่มาพร้อมกับ AGP 9.2.0 จะสแกน bytecode ที่คอมไพล์แล้วเพื่อหาแพทเทิร์น AtomicFieldUpdater มาตรฐาน เมื่อพบแล้ว จะทำการแปลงข้อมูล 4 ขั้นตอน:

  1. Identify the field’s memory offset. R8 computes the exact location of the target field inside the object layout.
  2. ระบุตำแหน่งหน่วยความจำ (memory offset) ของฟิลด์ R8 จะคำนวณตำแหน่งที่แน่นอนของฟิลด์เป้าหมายภายในโครงสร้างออบเจกต์
  3. Drop the updater object. The reflective wrapper disappears, saving memory and eliminating class-loading work.
  4. ตัดออบเจกต์ updater ออก ตัว wrapper แบบ reflective จะหายไป ช่วยประหยัดหน่วยความจำและลดภาระในการโหลดคลาส
  5. Insert a direct sun.misc.Unsafe call. This low-level API writes to the field using a single atomic hardware instruction.
  6. แทรกการเรียกใช้ sun.misc.Unsafe โดยตรง API ระดับต่ำนี้จะเขียนข้อมูลลงในฟิลด์โดยใช้คำสั่งฮาร์ดแวร์แบบ atomic เพียงคำสั่งเดียว
  7. Replace every updater call with the new unsafe instruction, allowing the JIT compiler to inline the operation.
  8. แทนที่การเรียกใช้งาน updater ทุกจุด ด้วยคำสั่ง unsafe ใหม่ ซึ่งช่วยให้ JIT compiler สามารถทำ inline การทำงานได้

The net effect is that the CPU no longer has to perform a reflective lookup or runtime checks; it executes the atomic instruction directly. From a developer’s perspective the change is invisible – the coroutine API behaves the same – but under the hood the code runs at “metal-level” speed.

ผลลัพธ์สุทธิคือ CPU ไม่ต้องทำการค้นหาแบบ reflective หรือตรวจสอบในขณะรันไทม์อีกต่อไป แต่จะรันคำสั่ง atomic โดยตรง ในมุมมองของนักพัฒนา การเปลี่ยนแปลงนี้จะไม่เห็นความแตกต่าง — coroutine API ยังคงทำงานเหมือนเดิม — แต่ภายใต้ระบบ โค้ดจะทำงานด้วยความเร็วในระดับ "metal-level"

Measurable gains

ผลลัพธ์ที่วัดได้

Benchmarks on a typical Android device show the following speed-ups after building with AGP 9.2.0, R8 enabled and minification turned on:

ผลการทดสอบ (Benchmarks) บนอุปกรณ์ Android ทั่วไปแสดงให้เห็นถึงความเร็วที่เพิ่มขึ้นดังนี้ หลังจาก Build ด้วย AGP 9.2.0, เปิดใช้งาน R8 และเปิดการทำ minification:

  • Channel send/receive: 2.01 × faster
  • Channel send/receive: เร็วขึ้น 2.01 ×
  • Mutex lock/unlock: 1.90 × faster
  • Mutex lock/unlock: เร็วขึ้น 1.90 ×
  • StateFlow updates: 2.02 × faster
  • StateFlow updates: เร็วขึ้น 2.02 ×
  • Coroutine dispatch: 1.68 × faster
  • Coroutine dispatch: เร็วขึ้น 1.68 ×

These numbers translate into tangible user-experience improvements. A cold launch that spent a fraction of a second waiting on coroutine synchronisation now finishes sooner, giving the UI thread more headroom to render the first frame. Less CPU contention also lets the processor return to sleep faster, which can improve battery life.

ตัวเลขเหล่านี้เปลี่ยนเป็นประสบการณ์ผู้ใช้ที่จับต้องได้ การเปิดแอปแบบ cold launch ที่เคยต้องรอการซิงโครไนซ์ของ coroutine เพียงเสี้ยววินาที จะเสร็จสิ้นเร็วขึ้น ทำให้ UI thread มีพื้นที่ว่าง (headroom) มากขึ้นในการเรนเดอร์เฟรมแรก การลดการแย่งชิงทรัพยากร CPU (CPU contention) ยังช่วยให้โปรเซสเซอร์กลับเข้าสู่โหมด sleep ได้เร็วขึ้น ซึ่งช่วยปรับปรุงอายุการใช้งานแบตเตอรี่ได้

How to reap the benefit

วิธีรับประโยชน์

No code changes are required. To activate the rewrite you need:

ไม่จำเป็นต้องแก้ไขโค้ดใดๆ ในการเปิดใช้งานการเขียนโค้ดใหม่นี้ คุณต้องมี:

  • AGP 9.2.0 or newer – the version that contains the updated R8.
  • AGP 9.2.0 หรือใหม่กว่า – เวอร์ชันที่มี R8 เวอร์ชันอัปเดต
  • R8 – automatically used when you build with the above AGP.
  • R8 – จะถูกใช้งานโดยอัตโนมัติเมื่อคุณ Build ด้วย AGP ข้างต้น
  • Kotlin Coroutines 1.8.0+ – the library version that ships with the AtomicFieldUpdater pattern the optimizer expects.
  • Kotlin Coroutines 1.8.0+ – เวอร์ชันของไลบรารีที่มาพร้อมกับแพทเทิร์น AtomicFieldUpdater ตามที่ตัว optimizer คาดหวัง
  • isMinifyEnabled = true in your release build type – R8 only runs when minification is on.
  • isMinifyEnabled = true ใน release build type ของคุณ – R8 จะทำงานเมื่อเปิดการทำ minification เท่านั้น

The only extra step is to audit your ProGuard (or R8) rules. Broad -keep directives that preserve volatile fields or the updater classes themselves block the rewrite. Make sure the rules allow R8 to modify those fields; otherwise the optimizer will fall back to the original reflective implementation.

ขั้นตอนเพิ่มเติมเพียงอย่างเดียวคือการตรวจสอบกฎ ProGuard (หรือ R8) ของคุณ คำสั่ง -keep ที่กว้างเกินไปซึ่งรักษาฟิลด์แบบ volatile หรือตัวคลาส updater เอง จะขัดขวางการเขียนโค้ดใหม่ ตรวจสอบให้แน่ใจว่ากฎของคุณอนุญาตให้ R8 แก้ไขฟิลด์เหล่านั้นได้ มิฉะนั้น optimizer จะกลับไปใช้การทำงานแบบ reflective ดั้งเดิม

You can verify the transformation with Android Studio’s APK Analyzer. Open the compiled APK, locate a coroutine support class such as JobSupport, and inspect the decompiled bytecode. If the rewrite succeeded, the static updater fields will be absent and you’ll see direct calls to Unsafe instead.

คุณสามารถตรวจสอบการแปลงข้อมูลได้ด้วย APK Analyzer ใน Android Studio โดยเปิด APK ที่คอมไพล์แล้ว ค้นหาคลาสสนับสนุน coroutine เช่น JobSupport และตรวจสอบ bytecode ที่ถูก decompiled หากการเขียนโค้ดใหม่สำเร็จ ฟิลด์ updater แบบ static จะหายไป และคุณจะเห็นการเรียกใช้ Unsafe โดยตรงแทน

Caveats and counter-points

ข้อควรระวังและมุมมองที่ต่างออกไป

The optimization hinges on two conditions that not every project meets:

การเพิ่มประสิทธิภาพนี้ขึ้นอยู่กับเงื่อนไขสองประการที่อาจไม่ได้มีในทุกโปรเจกต์:

  1. ต้องเปิดใช้งาน Minification Build แบบ Debug หรือ Build แบบ Release ที่ปิด Minification ไว้เพื่อความสะดวกในการดีบั๊ก จะไม่ได้รับประโยชน์นี้
  2. กฎของ ProGuard ต้องมีความยืดหยุ่น โปรเจกต์ที่มีรูปแบบ -keep ที่เข้มงวดเกินไปสำหรับส่วนประกอบภายในของ coroutine อาจจำเป็นต้องผ่อนปรนกฎเหล่านั้น ซึ่งอาจทำให้คลาสภายในเสี่ยงต่อบั๊กที่เกี่ยวข้องกับการทำ shrinking หากไม่มีการทดสอบอย่างระมัดระวัง

การทดสอบกับอุปกรณ์กลุ่มเป้าหมายที่หลากหลายยังคงเป็นแนวทางปฏิบัติที่ดี

สิ่งที่ควรสังเกตต่อไป

การ rewrite นี้แสดงให้เห็นว่าการแปลง bytecode ในช่วง build time สามารถดึงประสิทธิภาพที่ซ่อนอยู่ภายใต้การทำงานแบบ abstraction ของภาษาออกมาได้อย่างไร ควรพิจารณาทำ profiling กับโค้ดของคุณที่มีการใช้งาน coroutine อย่างหนัก เพื่อยืนยันประสิทธิภาพที่เพิ่มขึ้นใน workload เฉพาะของคุณ

สรุปสาระสำคัญ: การอัปเกรดเป็น AGP 9.2.0 และการเปิดใช้งาน minification ของ R8 ช่วยให้แอปพลิเคชัน Kotlin ที่มีการใช้งาน coroutine อย่างหนักมีความเร็วในการทำ synchronization ที่สำคัญเพิ่มขึ้นเกือบสองเท่า โดยไม่ต้องแก้ไขโค้ดต้นฉบับเลย—หากการตั้งค่า build อนุญาตให้ optimizer ทำงานได้อย่างเต็มที่