Computer Vision2019intermediate13 min read

EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

EfficientNet: إعادة التفكير في توسيع نماذج الشبكات الالتفافية

Tan, M. · Le, Q.V. — ICML

The problem

Convolutional neural networks are typically developed at a fixed computational budget, then scaled up when more resources become available. The standard approach was to scale only one dimension — make the network deeper (more layers), wider (more channels), or feed it higher- images. But each of these approaches hits diminishing returns quickly. Deeper networks suffer from vanishing gradients. Wider networks plateau in . Higher resolution alone cannot capture richer patterns. There was no principled method to scale all three dimensions together.

The contribution

A method that uniformly scales network depth, width, and input resolution using a single compound coefficient φ. The key insight is that these three dimensions are interdependent: higher-resolution images need deeper networks to capture finer patterns and wider layers to capture more per-pixel features. The authors used to design a baseline network (EfficientNet-B0, similar to MnasNet), then applied compound scaling to produce a family of models (B0–B7). EfficientNet-B7 achieved 84.3% on — state-of-the-art — while being 8.4× smaller and 6.1× faster than the best existing models.

The impact

EfficientNet proved that intelligent scaling outperforms brute-force model enlargement. Its compound scaling principle influenced every subsequent efficient architecture design. The EfficientNet family became the default backbone for across computer vision, replacing ResNet in many production systems. The work also highlighted neural architecture search as a practical tool, not just a research curiosity.

Imagine you are baking a cake and want to make it bigger. You could make it taller by adding more layers, wider by using a larger pan, or use finer ingredients for richer texture. Most bakers only do one of these at a time — and the result is lopsided: a very tall, thin cake that topples, or a wide, flat one with no depth.

This paper discovered the recipe: scale height, width, and ingredient quality together, in a fixed ratio. The result is a perfectly proportioned cake that is bigger and better — using less batter than the lopsided approaches.

The problem: scaling one dimension at a time hits a wall

Before EfficientNet, making a better usually meant making it bigger in one direction. ResNet scaled depth — going from 18 to 152 layers. WideResNet scaled width — doubling or quadrupling the channels. Some approaches scaled input resolution — feeding larger images (from 224×224 to 480×480 or beyond).

Each approach works... up to a point. Deeper networks can extract more abstract features, but eventually they suffer from vanishing gradients and difficulty, even with residual connections. Wider networks capture more per-layer features, but they plateau quickly — doubling the width gives diminishing accuracy gains. Higher resolution gives the network more pixels to work with, but without deeper or wider layers to process them, the extra detail goes to waste.

The key observation: each dimension has diminishing returns when scaled alone, but they are not independent. A higher-resolution image has more fine-grained patterns, so it needs a deeper network to capture them. And more layers processing more patterns need more channels to store what they find. The dimensions want to grow together.

Open in Lab
Toggle depth, width, and resolution independently to see diminishing returns — then enable compound scaling to see the balanced improvement.
The demo wakes as you arrive…

Compound scaling: one knob to rule three dimensions

The central contribution of this paper is the compound scaling method. Instead of scaling depth, width, or resolution independently, scale all three together using a single compound coefficient ϕ\phi.

The intuition is this: ϕ\phi is a budget knob. When you turn it up, you are saying "I have more compute to spend." The method then distributes that budget across all three dimensions in a fixed ratio, determined once by a small on the baseline model.

Why a fixed ratio? Because the authors empirically found that the optimal balance between depth, width, and resolution stays roughly constant regardless of scale. A model that needs 2×2\times more should get a bit deeper, a bit wider, and a bit more resolution — in the same proportions. This is a : the recipe for a good small model is the same recipe for a good large model, just multiplied.

Before seeing the formula, let us understand what each variable controls. The network has a baseline depth dd (number of layers), baseline width ww (number of channels), and baseline input resolution rr. The compound coefficient ϕ\phi controls how much to scale all three. The constants α\alpha, β\beta, and γ\gamma determine how ϕ\phi is distributed — how much of the extra budget goes to depth vs width vs resolution.

The constraint α⋅β2⋅γ2≈2\alpha \cdot \beta^2 \cdot \gamma^2 \approx 2 ensures that every time you increase ϕ\phi by 1, the total FLOPs roughly double. Why the squares? Because FLOPs scale linearly with depth but quadratically with width and resolution — doubling the channels quadruples the computation in each , and doubling the resolution quadruples the number of pixels each layer must process.

d=αϕ,w=βϕ,r=γϕs.t.  α⋅β2⋅γ2≈2,  α≥1,  β≥1,  γ≥1d = \alpha^{\phi}, \quad w = \beta^{\phi}, \quad r = \gamma^{\phi} \quad \text{s.t.} \; \alpha \cdot \beta^2 \cdot \gamma^2 \approx 2, \; \alpha \ge 1, \; \beta \ge 1, \; \gamma \ge 1
Compound scaling — depth, width, and resolution grow together with φ — The user-specified coefficient ϕ\phi controls the total compute budget: ϕ=1\phi=1 roughly doubles FLOPs over the baseline. The constants α=1.2\alpha=1.2, β=1.1\beta=1.1, γ=1.15\gamma=1.15 (found by grid search on B0) determine how that budget is split. Increasing ϕ\phi from 1 to 7 produces EfficientNet-B1 through B7.
Open in Lab
Adjust φ to see how depth, width, and resolution scale together. Compare with single-dimension scaling to see the efficiency advantage.
The demo wakes as you arrive…

The constraint α⋅β2⋅γ2≈2\alpha \cdot \beta^2 \cdot \gamma^2 \approx 2 deserves a closer look. The total FLOPs of a convolutional network are approximately proportional to d⋅w2⋅r2d \cdot w^2 \cdot r^2. Substituting the scaling formulas: total FLOPs ∝αϕ⋅(βϕ)2⋅(γϕ)2=(α⋅β2⋅γ2)ϕ\propto \alpha^\phi \cdot (\beta^\phi)^2 \cdot (\gamma^\phi)^2 = (\alpha \cdot \beta^2 \cdot \gamma^2)^\phi. By setting α⋅β2⋅γ2=2\alpha \cdot \beta^2 \cdot \gamma^2 = 2, we get FLOPs ∝2ϕ\propto 2^\phi. So ϕ\phi is simply the log-base-2 of the compute multiplier.

FLOPs∝(α⋅β2⋅γ2)ϕ=2ϕ\text{FLOPs} \propto (\alpha \cdot \beta^2 \cdot \gamma^2)^{\phi} = 2^{\phi}
FLOPs grow as 2^φ — φ = 1 means 2× compute, φ = 2 means 4×, etc. — This is why the constraint exists: it makes ϕ\phi a clean, interpretable budget control. Each integer step doubles compute. The constants α\alpha, β\beta, γ\gamma only need to be found once — they define the shape of the scaling curve.

The foundation: EfficientNet-B0 and MBConv blocks

Compound scaling only works if the baseline network is good. A poor starting architecture, no matter how cleverly scaled, will stay mediocre. So the authors used neural architecture search — the same multi-objective search from MnasNet — to find an efficient baseline optimized for both accuracy and FLOPs.

The result is EfficientNet-B0: a 5.3 million parameter network that achieves 77.1% top-1 accuracy on ImageNet with only 0.39 billion FLOPs. Its building blocks are blocks — Mobile inverted Convolution blocks, first introduced in MobileNetV2. Each MBConv block is like a small factory with three stages.

Stage 1 — Expand: A 1×1 convolution expands the number of channels by a factor of 6 (MBConv6) or keeps them the same (MBConv1). Think of this as widening the highway so more information can travel in parallel.

Stage 2 — Process: A applies one filter per (3×3 or 5×5), followed by and the . This is where the spatial pattern recognition happens — cheaply, because each filter operates on a single channel rather than across all channels.

Stage 3 — Compress + Recalibrate: A module learns which channels are most important and rescales them accordingly. Then a 1×1 convolution projects back to a smaller number of channels. If the input and output dimensions match, a residual adds the input directly to the output.

Open in Lab
Follow data through an MBConv block: expansion, depthwise convolution, squeeze-and-excitation, projection, and skip connection.
The demo wakes as you arrive…

Channel attention: letting the network focus

Not all channels in a are equally important. Some capture edges, others capture textures, others capture backgrounds. The -and- module is like a panel of judges that scores each channel's importance and amplifies the useful ones while suppressing the noisy ones.

Squeeze: compresses each channel's spatial map (H×W) into a single number — a summary of "how active is this channel overall?" This produces a vector of length CC (one number per channel).

Excitation: Two small fully-connected layers (with a reduction ratio of 0.25) process this vector and output a weight between 0 and 1 for each channel, using a sigmoid activation. These weights rescale the original feature map channel by channel.

Open in Lab
Watch how squeeze-and-excitation compresses spatial info, scores each channel, and rescales the feature map.
The demo wakes as you arrive…

The EfficientNet family: from B0 to B7

With the baseline B0 and the compound scaling formula in hand, producing the rest of the family is mechanical. Fix α=1.2\alpha=1.2, β=1.1\beta=1.1, γ=1.15\gamma=1.15 (found by grid search with ϕ=1\phi=1 on B0), then increase ϕ\phi to get larger models.

EfficientNet-B0 starts at 224×224 input resolution with 5.3M parameters. By ϕ=7\phi=7 (EfficientNet-B7), the network takes 600×600 images, has 66M parameters, and reaches 84.3% top-1 accuracy on ImageNet. For comparison, GPipe (the previous state-of-the-art) needed 557M parameters — 8.4× more — and was 6.1× slower at .

The models also transfer well. On CIFAR-100, Flowers, Stanford Cars, and other downstream datasets, EfficientNet variants achieved state-of-the-art accuracy with an order of magnitude fewer parameters than competing architectures.

Open in Lab
Compare EfficientNet B0–B7 against other architectures on accuracy vs FLOPs. Each point is a model — EfficientNets dominate the Pareto frontier.
The demo wakes as you arrive…

Why compound scaling works: the interaction effect

The paper provides empirical evidence for why compound scaling outperforms single-dimension scaling. The authors scaled a baseline network along each dimension independently and then combined them.

Scaling only depth (more layers) improves the network's ability to capture complex features, but the accuracy gain saturates. Very deep networks suffer from optimization difficulty — the training signal weakens as it travels through many layers.

Scaling only width (more channels) captures more fine-grained features per layer, but wide-and-shallow networks struggle to learn higher-level abstractions. The extra channels carry more information but lack the depth to combine it.

Scaling only resolution (larger images) provides more detail for the network to examine, but without extra depth and width, the network cannot fully exploit the information. It is like zooming in with a low-resolution camera — more area, same blur.

Scaling all three together avoids all these bottlenecks. Deeper layers can process the finer patterns visible in high-resolution inputs. Wider channels can store the richer features that deeper layers extract. Each dimension amplifies the others.

Code: EfficientNet-B0 in practice

Compound scaling coefficients for EfficientNet B0–B7python

Simplified to show the idea — not the real implementation.

# Compound scaling: once alpha, beta, gamma are found for B0, # higher models just increase phi. # alpha * beta**2 * gamma**2 ≈ 2  =>  1.2 * 1.1**2 * 1.15**2 ≈ 2.0
import math
alpha, beta, gamma = 1.2, 1.1, 1.15
# Verify the constraint print(f"alpha * beta^2 * gamma^2 = {alpha * beta**2 * gamma**2:.3f}")  # ≈ 2.0
# Generate the family for phi in range(8):                             # B0 (phi=0) to B7 (phi=7)
    d = alpha ** phi                             # depth multiplier
    w = beta  ** phi                             # width multiplier
    r = int(gamma ** phi * 224)                  # input resolution (base 224)
    flops_mult = 2 ** phi                        # approx FLOPs multiplier
    print(f"B{phi}: depth×{d:.2f}  width×{w:.2f}  "
          f"res={r}  FLOPs≈{flops_mult}×")

Results: dominating the accuracy-efficiency tradeoff

EfficientNet-B7 achieved 84.3% top-1 accuracy on ImageNet, matching GPipe — but with 8.4× fewer parameters (66M vs 557M) and 6.1× faster inference. EfficientNet-B0 alone, at 77.1% with 5.3M parameters, already matched the accuracy of ResNet-50 (76.0%) with fewer parameters and FLOPs.

On transfer learning benchmarks, EfficientNets dominated. On CIFAR-100, B7 achieved 91.7% accuracy. On Flowers, 98.8%. On Stanford Cars, comparable results — all with an order of magnitude fewer parameters than alternatives like SENet and NASNet-A.

The compound scaling method also proved general: when applied to MobileNetV1 and MobileNetV2 (not just the NAS-found baseline), it improved their accuracy by 1.4–2.0% over single-dimension scaling. This showed that the principle works across architectures, not just for one specific network.

Limitations and later insights

The compound scaling method assumes a fixed ratio between dimensions at all scales. Later work (EfficientNetV2, 2021) showed that this assumption breaks at very large scales — bigger models may benefit from relatively less depth and more width than the original formula suggests. The fixed constraint is a simplification of a more complex landscape.

Training large EfficientNets is also slow due to the high input resolution. EfficientNet-B7 trains on 600×600 images, which is expensive in GPU memory and training time. EfficientNetV2 addressed this with progressive resizing — starting with smaller images and gradually increasing resolution during training.

The reliance on neural architecture search means the baseline cannot be easily reproduced or intuitively understood — it is a black-box design. This contrasts with manually designed architectures like ResNet or VGG, where the design principles are transparent.

Legacy: from compound scaling to modern efficient networks

  1. 2017

    MobileNetV1 & depthwise separable convolutions

    Introduced depthwise separable convolutions for mobile deployment. Dramatically reduced computation per layer but limited accuracy at larger scales.

  2. 2018

    MobileNetV2 & the inverted residual block (MBConv)

    Introduced the inverted residual block — expand channels, process with depthwise convolution, then compress. This became the building block of EfficientNet.

  3. 2018

    SENet — Squeeze-and-Excitation Networks

    Won ImageNet 2017. Showed that channel attention — learning which channels matter — adds accuracy with minimal extra cost. EfficientNet integrates SE into every MBConv.

  4. 2019

    EfficientNet (this paper)

    United compound scaling + NAS baseline + MBConv + SE blocks. EfficientNet-B7 set new ImageNet SOTA at 84.3% with 8.4× fewer parameters than GPipe.

  5. 2021

    EfficientNetV2 — faster training + adaptive scaling

    Added Fused-MBConv blocks and progressive resizing. Showed that optimal scaling ratios change with model size, refining the compound scaling principle.

  6. 2022

    ConvNeXt — revisiting pure ConvNets

    Modernized ResNet with Transformer-era training recipes. Showed that well-designed ConvNets can compete with Vision Transformers, inheriting the scaling lessons from EfficientNet.

EfficientNet's deepest contribution is not any single model — it is the insight that scaling is a design problem, not just an engineering one. Before this paper, "make it bigger" was the default strategy. After it, every architecture team asks: how should we scale? The compound scaling principle — that depth, width, and resolution are interdependent and should grow together — has become foundational knowledge in neural network design.

CitationTan, Le. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. ICML, 2019.

Terms in this paper