Computer Vision2019intermediate13 min read
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
EfficientNet: إعادة التفكير في توسيع نماذج الشبكات الالتفافية
Tan, M. · Le, Q.V. — ICML
The problem
Convolutional neural networks are typically developed at a fixed computational budget, then scaled up when more resources become available. The standard approach was to scale only one dimension — make the network deeper (more layers), wider (more channels), or feed it higher- images. But each of these approaches hits diminishing returns quickly. Deeper networks suffer from vanishing gradients. Wider networks plateau in . Higher resolution alone cannot capture richer patterns. There was no principled method to scale all three dimensions together.
The contribution
A method that uniformly scales network depth, width, and input resolution using a single compound coefficient φ. The key insight is that these three dimensions are interdependent: higher-resolution images need deeper networks to capture finer patterns and wider layers to capture more per-pixel features. The authors used to design a baseline network (EfficientNet-B0, similar to MnasNet), then applied compound scaling to produce a family of models (B0–B7). EfficientNet-B7 achieved 84.3% on — state-of-the-art — while being 8.4× smaller and 6.1× faster than the best existing models.
The impact
EfficientNet proved that intelligent scaling outperforms brute-force model enlargement. Its compound scaling principle influenced every subsequent efficient architecture design. The EfficientNet family became the default backbone for across computer vision, replacing ResNet in many production systems. The work also highlighted neural architecture search as a practical tool, not just a research curiosity.
Imagine you are baking a cake and want to make it bigger. You could make it taller by adding more layers, wider by using a larger pan, or use finer ingredients for richer texture. Most bakers only do one of these at a time — and the result is lopsided: a very tall, thin cake that topples, or a wide, flat one with no depth.
This paper discovered the recipe: scale height, width, and ingredient quality together, in a fixed ratio. The result is a perfectly proportioned cake that is bigger and better — using less batter than the lopsided approaches.
The problem: scaling one dimension at a time hits a wall
Before EfficientNet, making a better usually meant making it bigger in one direction. ResNet scaled depth — going from 18 to 152 layers. WideResNet scaled width — doubling or quadrupling the channels. Some approaches scaled input resolution — feeding larger images (from 224×224 to 480×480 or beyond).
Each approach works... up to a point. Deeper networks can extract more abstract features, but eventually they suffer from vanishing gradients and difficulty, even with residual connections. Wider networks capture more per-layer features, but they plateau quickly — doubling the width gives diminishing accuracy gains. Higher resolution gives the network more pixels to work with, but without deeper or wider layers to process them, the extra detail goes to waste.
The key observation: each dimension has diminishing returns when scaled alone, but they are not independent. A higher-resolution image has more fine-grained patterns, so it needs a deeper network to capture them. And more layers processing more patterns need more channels to store what they find. The dimensions want to grow together.
Compound scaling: one knob to rule three dimensions
The central contribution of this paper is the compound scaling method. Instead of scaling depth, width, or resolution independently, scale all three together using a single compound coefficient .
The intuition is this: is a budget knob. When you turn it up, you are saying "I have more compute to spend." The method then distributes that budget across all three dimensions in a fixed ratio, determined once by a small on the baseline model.
Why a fixed ratio? Because the authors empirically found that the optimal balance between depth, width, and resolution stays roughly constant regardless of scale. A model that needs more should get a bit deeper, a bit wider, and a bit more resolution — in the same proportions. This is a : the recipe for a good small model is the same recipe for a good large model, just multiplied.
Before seeing the formula, let us understand what each variable controls. The network has a baseline depth (number of layers), baseline width (number of channels), and baseline input resolution . The compound coefficient controls how much to scale all three. The constants , , and determine how is distributed — how much of the extra budget goes to depth vs width vs resolution.
The constraint ensures that every time you increase by 1, the total FLOPs roughly double. Why the squares? Because FLOPs scale linearly with depth but quadratically with width and resolution — doubling the channels quadruples the computation in each , and doubling the resolution quadruples the number of pixels each layer must process.
The constraint deserves a closer look. The total FLOPs of a convolutional network are approximately proportional to . Substituting the scaling formulas: total FLOPs . By setting , we get FLOPs . So is simply the log-base-2 of the compute multiplier.
The foundation: EfficientNet-B0 and MBConv blocks
Compound scaling only works if the baseline network is good. A poor starting architecture, no matter how cleverly scaled, will stay mediocre. So the authors used neural architecture search — the same multi-objective search from MnasNet — to find an efficient baseline optimized for both accuracy and FLOPs.
The result is EfficientNet-B0: a 5.3 million parameter network that achieves 77.1% top-1 accuracy on ImageNet with only 0.39 billion FLOPs. Its building blocks are blocks — Mobile inverted Convolution blocks, first introduced in MobileNetV2. Each MBConv block is like a small factory with three stages.
Stage 1 — Expand: A 1×1 convolution expands the number of channels by a factor of 6 (MBConv6) or keeps them the same (MBConv1). Think of this as widening the highway so more information can travel in parallel.
Stage 2 — Process: A applies one filter per (3×3 or 5×5), followed by and the . This is where the spatial pattern recognition happens — cheaply, because each filter operates on a single channel rather than across all channels.
Stage 3 — Compress + Recalibrate: A module learns which channels are most important and rescales them accordingly. Then a 1×1 convolution projects back to a smaller number of channels. If the input and output dimensions match, a residual adds the input directly to the output.
Channel attention: letting the network focus
Not all channels in a are equally important. Some capture edges, others capture textures, others capture backgrounds. The -and- module is like a panel of judges that scores each channel's importance and amplifies the useful ones while suppressing the noisy ones.
Squeeze: compresses each channel's spatial map (H×W) into a single number — a summary of "how active is this channel overall?" This produces a vector of length (one number per channel).
Excitation: Two small fully-connected layers (with a reduction ratio of 0.25) process this vector and output a weight between 0 and 1 for each channel, using a sigmoid activation. These weights rescale the original feature map channel by channel.
The EfficientNet family: from B0 to B7
With the baseline B0 and the compound scaling formula in hand, producing the rest of the family is mechanical. Fix , , (found by grid search with on B0), then increase to get larger models.
EfficientNet-B0 starts at 224×224 input resolution with 5.3M parameters. By (EfficientNet-B7), the network takes 600×600 images, has 66M parameters, and reaches 84.3% top-1 accuracy on ImageNet. For comparison, GPipe (the previous state-of-the-art) needed 557M parameters — 8.4× more — and was 6.1× slower at .
The models also transfer well. On CIFAR-100, Flowers, Stanford Cars, and other downstream datasets, EfficientNet variants achieved state-of-the-art accuracy with an order of magnitude fewer parameters than competing architectures.
Why compound scaling works: the interaction effect
The paper provides empirical evidence for why compound scaling outperforms single-dimension scaling. The authors scaled a baseline network along each dimension independently and then combined them.
Scaling only depth (more layers) improves the network's ability to capture complex features, but the accuracy gain saturates. Very deep networks suffer from optimization difficulty — the training signal weakens as it travels through many layers.
Scaling only width (more channels) captures more fine-grained features per layer, but wide-and-shallow networks struggle to learn higher-level abstractions. The extra channels carry more information but lack the depth to combine it.
Scaling only resolution (larger images) provides more detail for the network to examine, but without extra depth and width, the network cannot fully exploit the information. It is like zooming in with a low-resolution camera — more area, same blur.
Scaling all three together avoids all these bottlenecks. Deeper layers can process the finer patterns visible in high-resolution inputs. Wider channels can store the richer features that deeper layers extract. Each dimension amplifies the others.
Code: EfficientNet-B0 in practice
Simplified to show the idea — not the real implementation.
# Compound scaling: once alpha, beta, gamma are found for B0, # higher models just increase phi. # alpha * beta**2 * gamma**2 ≈ 2 => 1.2 * 1.1**2 * 1.15**2 ≈ 2.0
import math
alpha, beta, gamma = 1.2, 1.1, 1.15
# Verify the constraint print(f"alpha * beta^2 * gamma^2 = {alpha * beta**2 * gamma**2:.3f}") # ≈ 2.0
# Generate the family for phi in range(8): # B0 (phi=0) to B7 (phi=7)
d = alpha ** phi # depth multiplier
w = beta ** phi # width multiplier
r = int(gamma ** phi * 224) # input resolution (base 224)
flops_mult = 2 ** phi # approx FLOPs multiplier
print(f"B{phi}: depth×{d:.2f} width×{w:.2f} "
f"res={r} FLOPs≈{flops_mult}×")Results: dominating the accuracy-efficiency tradeoff
EfficientNet-B7 achieved 84.3% top-1 accuracy on ImageNet, matching GPipe — but with 8.4× fewer parameters (66M vs 557M) and 6.1× faster inference. EfficientNet-B0 alone, at 77.1% with 5.3M parameters, already matched the accuracy of ResNet-50 (76.0%) with fewer parameters and FLOPs.
On transfer learning benchmarks, EfficientNets dominated. On CIFAR-100, B7 achieved 91.7% accuracy. On Flowers, 98.8%. On Stanford Cars, comparable results — all with an order of magnitude fewer parameters than alternatives like SENet and NASNet-A.
The compound scaling method also proved general: when applied to MobileNetV1 and MobileNetV2 (not just the NAS-found baseline), it improved their accuracy by 1.4–2.0% over single-dimension scaling. This showed that the principle works across architectures, not just for one specific network.
Limitations and later insights
The compound scaling method assumes a fixed ratio between dimensions at all scales. Later work (EfficientNetV2, 2021) showed that this assumption breaks at very large scales — bigger models may benefit from relatively less depth and more width than the original formula suggests. The fixed constraint is a simplification of a more complex landscape.
Training large EfficientNets is also slow due to the high input resolution. EfficientNet-B7 trains on 600×600 images, which is expensive in GPU memory and training time. EfficientNetV2 addressed this with progressive resizing — starting with smaller images and gradually increasing resolution during training.
The reliance on neural architecture search means the baseline cannot be easily reproduced or intuitively understood — it is a black-box design. This contrasts with manually designed architectures like ResNet or VGG, where the design principles are transparent.
Legacy: from compound scaling to modern efficient networks
2017
MobileNetV1 & depthwise separable convolutions
Introduced depthwise separable convolutions for mobile deployment. Dramatically reduced computation per layer but limited accuracy at larger scales.
2018
MobileNetV2 & the inverted residual block (MBConv)
Introduced the inverted residual block — expand channels, process with depthwise convolution, then compress. This became the building block of EfficientNet.
2018
SENet — Squeeze-and-Excitation Networks
Won ImageNet 2017. Showed that channel attention — learning which channels matter — adds accuracy with minimal extra cost. EfficientNet integrates SE into every MBConv.
2019
EfficientNet (this paper)
United compound scaling + NAS baseline + MBConv + SE blocks. EfficientNet-B7 set new ImageNet SOTA at 84.3% with 8.4× fewer parameters than GPipe.
2021
EfficientNetV2 — faster training + adaptive scaling
Added Fused-MBConv blocks and progressive resizing. Showed that optimal scaling ratios change with model size, refining the compound scaling principle.
2022
ConvNeXt — revisiting pure ConvNets
Modernized ResNet with Transformer-era training recipes. Showed that well-designed ConvNets can compete with Vision Transformers, inheriting the scaling lessons from EfficientNet.
EfficientNet's deepest contribution is not any single model — it is the insight that scaling is a design problem, not just an engineering one. Before this paper, "make it bigger" was the default strategy. After it, every architecture team asks: how should we scale? The compound scaling principle — that depth, width, and resolution are interdependent and should grow together — has become foundational knowledge in neural network design.
CitationTan, Le. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. ICML, 2019.
Terms in this paper
- Compound Scalingالتوسيع المُركَّب
- Neural Architecture Searchالبحث التلقائي عن بنية الشبكة الأمثل
- Efficiencyالكفاءة
- Depthwise Separable Convolutionالالتفاف القابل للفصل بالعمق
- Squeeze-and-Excitationالضغط والإثارة
- MBConvكتلة الاختناق المعكوس المتنقّلة
- Inverted Residual Blockكتلة الاختناق المعكوس
- SwishSwish
- FLOPsالعمليات الحسابية العائمة
- ImageNetImageNet
- Transfer Learningنقل التعلم
- Batch Normalizationتسوية الدفعات الحسابية
- Resolutionدقة الصورة
- Scaling Lawقانون التحجيم
- Dropoutالإسقاط العشوائي للعصبونات