Meta says it has reached an industrial milestone in training its artificial intelligence models for advertising. In a post by its Meta AI Engineering team, titled “GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model”, the group says it has doubled the training efficiency of GEM, a foundation model used for advertising recommendations on Facebook and Instagram.
At first glance, the message is less spectacular than the launch of a new large language model or a chatbot demonstration. Yet it is central to the economics of AI at scale. Meta says it has reduced the computing requirements needed to train GEM while preserving the model's performance. For a company whose advertising systems operate across billions of interactions, the ability to deliver the same level of quality with fewer computing resources can weigh heavily on costs, development timelines and deployment speed.
The announcement also highlights a sometimes less visible dimension of competition among AI labs: operational efficiency. Progress does not rely solely on larger models, more data or ever-larger processor clusters. It also depends on how models are trained, the software that uses accelerators, the organization of distributed computing and the ability to avoid unused resources.
GEM, a model designed for the core of Meta's advertising business
Meta presents GEM as an advertising foundation model at the scale of large language models. It powers advertising recommendation systems for Facebook and Instagram, the two long-standing services at the center of the group's commercial activity. This detail is essential: it is not a general-purpose model presented to the public as a conversational assistant, but an infrastructure component designed to improve recommendation decisions in a complex advertising environment.
In digital advertising, a recommendation does not simply consist of selecting an ad. Systems must process a multitude of signals and produce predictions used, among other things, to assess an ad's potential relevance in a given context. Platforms then relate these predictions to delivery constraints, advertisers' objectives and auction mechanisms. In the elements highlighted by this announcement, Meta does not detail the full range of specific uses covered by GEM or the associated commercial parameters. The group nevertheless says that the model serves advertising recommendations on Facebook and Instagram.
This direction is part of Meta's long technological history. Facebook, created in 2004, gradually built an advertising business based on personalization and measurement. Instagram, acquired by Facebook in 2012, became another pillar of this ecosystem. The group adopted the name Meta in 2021, but its recommendation and advertising infrastructures remain decisive to its business model. News feeds, video formats, recommended content and ads have long relied on algorithmic ranking systems.
The difference highlighted with GEM lies in its scale and its foundation-model logic. In AI terminology, a foundation model generally refers to a model trained at large scale that can be adapted to several tasks. In GEM's case, Meta emphasizes an “LLM-scale” approach, meaning a scale comparable, in terms of the order of magnitude of infrastructure and training methods, to that used for large language models. This comparison does not mean that GEM is a chatbot or that it is intended for the same uses as conversational models released by generative AI companies.
Meta AI Engineering's publication instead suggests that the group is applying some of the lessons learned from training large models to its advertising business. This is an important shift in how production AI is viewed. The most widely publicized models are often those that generate text, images, audio or code. But for major internet companies, recommendation and prediction models remain systems directly integrated into products used every day at very large scale.
Advertising is a particularly demanding field. A system must operate continuously, absorb massive volumes of data and make predictions within time frames compatible with ad delivery. It must also be trained and updated as content, behavior, campaigns and formats evolve. A model's performance is therefore inseparable from its training cost and operating cost. An efficiency gain is not merely a matter of technical optimization: it can change a company's ability to iterate more quickly on its systems.
The very title of the publication reveals this priority. Meta does not present GEM as an entirely new model or as an overhaul of Facebook's or Instagram's advertising formats. The announcement concerns training: “How Meta Doubled the Efficiency”. The issue is the model's production efficiency, at a time when AI expansion is running up against the availability of accelerators, the power consumption of data centers and the growing cost of computing infrastructure.
Meta claims to have doubled training efficiency
The main fact communicated by Meta is clear: the company says it has doubled GEM's training efficiency. In other words, according to the group, training this advertising foundation model can achieve the same level of results with a reduced amount of compute, while maintaining model performance. Meta presents this improvement as progress achieved at the scale of a system comparable, in terms of its training constraints, to large language models.
The word “efficiency” deserves to be read carefully. In the context of the announcement, it does not necessarily refer to a uniform reduction in all costs associated with an advertising system. Meta refers to training efficiency and computing requirements. Without additional data, the publication should therefore not be interpreted as announcing an exactly equivalent reduction in all advertising delivery, storage, network or inference expenses. The claimed scope concerns the model-training stage.
The distinction matters because training and inference are two different industrial realities. Training consists of adjusting a model's parameters from data and generally requires intensive computation on specialized infrastructure. Inference refers to the moment when the model is used to produce a prediction or recommendation. Both are essential, but they do not necessarily face the same technical constraints. Meta's announcement explicitly concerns the former.
The group also says that this gain did not come at the cost of lower performance. This is a crucial point. Computing requirements can be reduced by decreasing a model's size, shortening training or limiting the data processed, but these choices can also degrade the quality of the results. Meta's message is precisely that the optimizations made to GEM make it possible to avoid this trade-off, at least according to the internal evaluations to which the company refers.
The source publication does not provide, in the brief associated with this announcement, a financial amount, a number of processors, a training duration, a data volume or a detailed quality metric. It would therefore be risky to infer savings in euros, a quantified reduction in emissions or an acceleration expressed in days. The confirmed data point is the doubling of training efficiency claimed by Meta, along with the maintenance of model performance according to the company.
This caution does not diminish the result's potential significance. At the scale of a model used in the advertising recommendations of services such as Facebook and Instagram, halving the resources needed for training can make it easier to experiment with new configurations. It can also reduce pressure on scarce computing capacity. In modern AI, costs are not only linked to hardware purchases: they include the actual use of accelerators, transfers between machines, synchronization of computations and overall data-center operations.
Meta thus places GEM within a logic of industrial efficiency. When models reach a certain size, having high-performance hardware is no longer enough. That hardware must be used efficiently throughout the training cycle. Graphics processors or accelerators left idle, operations poorly distributed between machines or bottlenecks in data exchanges can sharply reduce the real efficiency of an infrastructure, even one that is very powerful on paper.
The choice to publish this work through Meta AI Engineering is also significant. Engineering blogs from major platforms often serve a dual purpose: they detail technical advances and signal to markets, researchers and candidates that the company has mastered the infrastructure required for modern models. Meta has communicated for several years about its work in AI, whether in fundamental research, open models or internal infrastructure. With GEM, the group explicitly links this effort to its advertising system.
For advertisers, this announcement alone does not change the interfaces, prices or tools available to them. Meta is not presenting a new commercial advertising product here. The issue is deeper but less directly visible: making the algorithmic layer that supports recommendations more efficient to train. This is precisely the kind of evolution that can have major consequences for a platform without being accompanied by a new button in Ads Manager or an immediately noticeable change in the applications.
The battle over models is also being fought in data centers
The recent rise of generative AI has popularized the idea that performance depends above all on ever-larger models. This view is partly correct: large models have required considerable volumes of compute and have pushed companies to invest in massive infrastructure. But it is incomplete. As deployments become more widespread, efficiency is becoming a differentiating factor just as important as a model's raw size.
Meta is not the only player to make optimization a priority. Companies developing large models, including OpenAI, Google, Microsoft, Anthropic and Amazon, regularly communicate about their systems' performance and infrastructure capabilities. However, public announcements take different forms. Some highlight new models and their capabilities in reasoning, programming or multimodal generation tasks. Others concern chips, cloud platforms or deployment tools.
GEM stands out because it is directly rooted in a long-standing Meta use case: advertising recommendation. The group is not merely presenting a laboratory improvement or an advance for an AI assistant. It describes an optimization tied to a model that supports delivery decisions on Facebook and Instagram. This proximity to the company's economic core gives the claimed gain particular value.
In this sector, infrastructure efficiency has become a matter of competitiveness. Access to computing accelerators is limited by production capacity and sustained global demand. The investments required to build and operate data centers are considerable. Under these conditions, a player able to obtain more results with the same compute budget can free up resources for other projects, reduce the time needed for experiments or improve the predictability of deployments.
However, a gain announced for GEM should not be turned into a general ranking of AI companies. Models, data, objectives and infrastructures differ greatly from one player to another. An advertising system is not evaluated like a conversational model. The quality metrics used in recommendation are not those used to measure an assistant's ability to write, answer questions or produce code. The more relevant comparison concerns strategic direction: all major players have an interest in extracting more value from every unit of compute.
Meta's communication also highlights a recurring tension in the industry. Companies want to demonstrate their ability to deploy increasingly sophisticated models, but they must simultaneously justify rising infrastructure spending. The profitability of AI uses will depend heavily on this equation. A highly capable model that is too costly to train or operate may remain difficult to generalize. Conversely, an efficiency improvement can turn a costly system into a technology that can be used more widely.
In this context, the vocabulary of “efficiency” does not refer only to marginal optimization. It describes a complete discipline, ranging from model design to the orchestration of computing tasks. Research into architectures, training methods and systems software has a direct economic effect. Advances of this kind are often less visible than a product release, but they can be more durable because they carry over into successive generations of models.
Meta has a structural advantage in this area: the company operates its own services at large scale and can integrate internal advances into systems that are already widely used. This does not mean that every optimization will automatically translate into a measurable improvement for users or advertisers. But it does mean that the company has an operational environment where training work can have an immediate application.
The notion of a foundation model further reinforces this interest. When a base can be reused across several tasks or product developments, the cost of its initial training takes on a strategic dimension. Reducing that cost, or making it more efficient, can increase a company's ability to update the model, experiment with variants and improve the systems dependent on that base. Meta does not detail all of GEM's applications, but its positioning as a foundation model suggests precisely an ambition of pooling.
Indirect consequences for advertisers and the French-speaking market
For French brands, agencies and companies that use Facebook and Instagram to communicate, the announcement does not come with a quantified promise regarding campaign results. Meta does not say that doubling GEM's training efficiency will mechanically lead to doubled advertising performance, lower acquisition costs or a stated increase in conversions. Such shortcuts would be unfounded.
The signal sent to the market is nevertheless important. Advertising recommendations increasingly rely on complex AI systems, and platforms are seeking to improve these systems without endlessly increasing computing expenditure. If Meta can train the models used in its ecosystem more efficiently, the company can potentially iterate more quickly on the technologies that support advertising delivery. This is a logical consequence of the claimed efficiency gain, not a commercial effect precisely quantified by Meta.
For French-speaking players, this development is a reminder that digital advertising is also an infrastructure issue. Advertisers mainly see dashboards, campaign objectives, audiences or creative assets. Behind the scenes, platforms operate models that learn from vast sets of signals and must be regularly trained. Architectural and optimization choices made by large technology companies therefore indirectly influence the environment in which campaigns are delivered.
In France, as in the rest of the European Union, this infrastructure is evolving within a demanding regulatory framework. The General Data Protection Regulation, or GDPR, has applied since 2018. The Digital Markets Act, which has gradually come into application for the very large platforms concerned, also governs certain digital practices. These texts do not directly address GEM's training efficiency, but they matter in the environment in which recommendation and advertising systems are designed and deployed.
More efficient training does not, by itself, answer questions of transparency, data protection, user control or competition. Computing optimization and regulatory compliance are two distinct dimensions. A platform can improve how it uses its computing resources while remaining subject to the same obligations regarding data processing and the presentation of its services. For European regulators, the issue remains assessing platforms' concrete practices, not only their technical performance.
The energy dimension should also be kept within its limits. Reducing the computing requirements associated with training a model can, in principle, reduce the resources needed for the same training operation. But Meta's announcement provides no quantified environmental assessment concerning GEM. It is therefore not possible to convert the announced efficiency gain into avoided electricity consumption, avoided emissions or a net impact on data centers. The overall effect would depend on many factors, including the number of training runs actually performed and changes in the scope of the models.
This caveat is particularly useful as European debates on AI increasingly incorporate resource issues. Large models are associated with substantial computing requirements. In this context, efficiency gains are likely to become an indicator monitored by companies and authorities, but they will need to be accompanied by transparent methodologies to enable robust comparisons. Meta's publication highlights an internal performance result; as it stands, it does not constitute complete environmental reporting.
For French companies developing their own AI systems, the announcement also illustrates a gap in resources. Few players can train models at the scale claimed by Meta for GEM. Yet the lesson is not limited to American giants: pipeline efficiency, the quality of hardware utilization and the reduction of unnecessary computation matter at every scale. In an environment where computing resources remain costly, optimization can be a more accessible lever than the race for size.
Cloud providers and European AI companies are also concerned. Demand for infrastructure suited to AI models is not limited to very large-scale training. It also concerns inference, data, development tools and governance. Meta's announcement confirms that technological advantage does not depend exclusively on access to hardware: it depends on the ability to make hardware, software and models work together under industrial conditions.
Toward more industrialized advertising AI, without an end to the race for compute
The prospect opened by GEM is that of increasingly industrialized AI. Meta highlights a training optimization, but its message extends beyond this model alone: when AI becomes central infrastructure for global products, efficiency must be considered from the design stage. Companies cannot indefinitely rely on adding raw capacity to solve all performance problems.
This evolution does not mean that the race for compute will stop. Major players continue to invest in data centers, accelerators and larger-scale models. Efficiency gains may even make it possible to reallocate resources to new training runs, new features or more ambitious models. There is a possible rebound effect: doing better with less compute per training run does not guarantee that the total compute used by a company will decline if the number of projects and the size of systems increase at the same time.
The exact significance of the announcement will therefore depend on how Meta uses this gain over time. The group says it is preserving GEM's performance despite lower computing requirements, which establishes a technical foundation. It remains to be seen whether the approach will be extended to other internal models, whether it will facilitate more frequent improvement cycles or whether it will influence future recommendation systems for Facebook and Instagram. The publication does not allow these developments to be asserted, but it makes these questions particularly relevant.
For the advertising sector, the most important change may remain barely visible. Advertisers will not judge GEM on its training efficiency as such, but on the stability of the tools, the relevance of recommendations, the clarity of controls and the results obtained in their own campaigns. Users, meanwhile, will continue to question the personalization of content and ads, regardless of the degree of optimization of the data centers that train the models.
Meta will therefore have to reconcile several imperatives: improving the industrial efficiency of its models, maintaining the quality of its recommendations, complying with the rules applicable in its markets and meeting transparency expectations. The announced doubling of efficiency for GEM does not resolve these balances, but it shows where a growing share of competition lies: in the ability to run AI at very large scale in an economically sustainable way.
In the coming years, the most decisive announcements may not always be those that present the most imposing model or the most viral demonstration. They may also come from infrastructure improvements such as the one claimed by Meta AI Engineering for GEM. As AI becomes integrated into digital services and recommendation mechanisms, the difference between a research innovation and an industrial advantage will increasingly depend on this simple but decisive question: how much compute is really needed to achieve the same level of performance?
Comments· 2 comments
The headline sounds impressive, but the article feels too reliant on Meta’s own framing. I would have liked more context on how “efficiency” is measured, what trade-offs might be involved, and whether advertisers or users would actually notice a meaningful difference.
That is a fair request, but a short report can still highlight the claimed technical milestone. If the efficiency gains hold up in practice without reducing recommendation quality, it seems reasonable to view it as potentially significant even before every external detail is available.