A standard convolutional layer is oddly egalitarian. It stacks dozens — sometimes hundreds — of feature detectors, then treats every single one with the same respect. A channel that has learned to spot diagonal edges gets the same vote as one that detects blue sky pixels. A channel firing on fur texture in a cat photo is broadcast forward with the same weight as one activated by background noise. That uniformity is the weakness. Not every channel matters equally for every image, and deep networks have historically lacked a mechanism to say so.
Squeeze-and-Excitation, introduced by Hu and colleagues, fixes this with a remarkably small addition. It attaches a learnable gating mechanism to any convolution block so the network can dynamically recalibrate how much each channel contributes, and it does this fresh for every input. The extra cost is trivial — roughly 2.5 percent more parameters when bolted onto ResNet-50 — but the gain in representational power is substantial enough that SE blocks helped win ILSVRC 2017 and now sit inside production architectures like MobileNetV3 and EfficientNet.
The idea is simple enough to build from scratch in three stages.
Squeeze: Giving the Network a Wide-Angle Lens
Convolution is local by design. A 3×3 kernel slides across the image and only ever sees its immediate neighborhood. Stack enough layers and the receptive field grows, yet no single operation looks at an entire feature map in one glance. That means a channel might be strongly activated across the whole spatial extent of a dog or a car, but the next layer never receives an explicit summary saying, “This detector fired everywhere.”
The squeeze step closes that gap with global average pooling. For each channel, every value in the H×W spatial grid is averaged down to a single number. If you have 512 channels, you now have 512 scalars. Each scalar encodes the global presence of whatever that channel has learned to detect — be it a wheel rim, a strip of grass, or a patch of skin tone. It is a channel-wise summary of the whole scene.
This matters because context changes meaning. A horizontal-line detector is useful in one image because it identifies a horizon, and useless in another because it is picking up noise in a forest canopy. Without spatial aggregation, the network lacks a clean signal to make that distinction.
Excite: A Bottleneck That Forces Cooperation
Once the squeeze step has produced a vector of global channel descriptors, the excite step learns what to do with them. This is a tiny two-layer MLP. It takes the vector, compresses it down to a bottleneck, and then expands it back to the original channel dimension.
Typical implementations use a reduction ratio of 16. If the input has 512 channels, the first linear layer projects the 512 scalars down to 32, a ReLU sits in between, and the second layer maps the 32 back up to 512. That compression is intentional. It works like a funnel: the network cannot simply pass the original values through unchanged. It must learn a compressed, shared representation of how channels relate to one another. The bottleneck forces cross-channel dependencies to emerge. The model might learn, for instance, that when a “bark texture” channel is strong, a “tree trunk” contour channel should also be emphasized, while a “water reflection” channel should be muted.
The final activation here is crucial, and it is a common place where intuition slips. Many people instinctively reach for softmax, because it feels like attention should be a probability distribution. Softmax would force all channels to compete against each other until their weights sum to one. If one channel goes up, others must come down. That is exactly wrong for this task. A photograph of a sunset may need both an orange-glow channel and a cloud-texture channel operating at full strength simultaneously. The gates should act like independent dimmer switches, not a zero-sum budget.
Sigmoid solves this. It squashes each scalar to a value between zero and one, but every channel is free to move independently. The model can turn any subset up, leave others neutral, and suppress the rest, all without artificial competition.
Scale: Apply the Gates and Move On
Langkah terakhir hampir terasa antiklimaks, yang justru merupakan bagian dari keanggunannya. Setiap feature map asli dikalikan dengan nilai gate yang diaktivasi oleh sigmoid yang sesuai. Jika gate mendekati satu, channel tersebut diteruskan tanpa perubahan. Jika gate mendekati nol, seluruh feature map tersebut diredam, memudar menuju keheningan. Hasilnya kemudian dimasukkan ke lapisan jaringan berikutnya.
Karena ini murni perkalian di seluruh peta spasial, hal ini menambah overhead komputasi yang minimal saat inferensi. Pekerjaan berat dilakukan oleh global pooling dan dua proyeksi kecil, yang nilainya dapat diabaikan dibandingkan dengan konvolusi 3×3 di sekitarnya.
Mengapa Desain Ini Bertahan
Di luar matematika yang bersih, Squeeze-and-Excitation bertahan karena ia menghormati batasan-batasan teknis.
Parameter yang murah. Menambahkan blok SE ke ResNet-50 hanya menambah sekitar 2,5 persen parameter. Lapisan MLP berukuran kecil; mereka tidak membuat model menjadi bengkak. Di era di mana peningkatan akurasi sering kali datang dari penggandaan ukuran model, SE menawarkan "free lunch" yang langka.
Modularitas. SE bukanlah arsitektur backbone baru yang menuntut Anda merancang ulang pipeline Anda. Ini adalah modul drop-in. Anda dapat menyisipkannya setelah blok konvolusi apa pun di ResNet, DenseNet, atau tumpukan kustom tanpa menulis ulang logika di sekitarnya. Fleksibilitas tersebut membuat adopsinya cepat, pertama di penelitian dan kemudian di jaringan yang dioptimalkan untuk seluler.
Perilaku adaptif input. Berbeda dengan bobot channel statis yang dipelajari selama pelatihan dan dibekukan, gate SE dihitung secara langsung (on the fly) untuk setiap forward pass. Berikan jaringan gambar langit yang datar dan tanpa fitur, maka channel yang relevan akan tetap diredam. Tunjukkan pemandangan jalanan yang sibuk dengan pejalan kaki, rambu-rambu, dan kendaraan, maka gate akan melakukan kalibrasi ulang untuk memperkuat detektor yang penting bagi gambar spesifik tersebut. Jaringan yang sama menyesuaikan sensitivitasnya sendiri adegan demi adegan.
Dari Pemenang Kompetisi Menjadi Implementasi Sehari-hari
Blok SE meraih posisi teratas di ILSVRC 2017, tetapi validasi sebenarnya datang kemudian. Mereka dimasukkan ke dalam MobileNetV3, di mana mereka membantu model ringan untuk memberikan performa yang melampaui kelasnya tanpa menguras baterai ponsel. Mereka juga muncul dalam resep compound scaling EfficientNet, membuktikan bahwa channel attention bukanlah kemewahan bagi server berat, melainkan alat praktis untuk desain yang efisien.
Jika Anda ingin melihat betapa sedikitnya kode yang sebenarnya dibutuhkan, demo langsung ini menguraikan blok tersebut baris demi baris:
- Live demo: https://dev48v.infy.uk/dl/day40-squeeze-excitation.html
- Full build walkthrough: https://dev.to/dev48v/squeeze-and-excitation-from-scratch-cheap-channel-attention-that-learns-which-feature-maps-matter-1hic
Dan jika Anda membangun bersama komunitas: https://t.me/GyaanSetuAi
Kesimpulan Utama
Model visi yang lebih baik tidak selalu membutuhkan tumpukan yang lebih dalam atau filter yang lebih lebar. Terkadang mereka hanya butuh momen untuk bertanya channel mana yang sebenarnya penting bagi gambar di hadapan mereka. Squeeze-and-Excitation memberikan momen itu hanya dengan rata-rata pooled, MLP yang terkompresi, dan gate sigmoid. Hasilnya adalah jaringan yang berhenti meneriakkan setiap fitur dengan volume yang sama dan sebaliknya belajar untuk berbisik, menekankan, atau membungkam setiap channel tepat saat adegan membutuhkannya.
