T0048/24 on Sufficiency of Disclosure for Machine Learning Inventions
One of the better known AI cases at the European Patent Office was Vesuvius vs. Refractory (T 1669/21), in which the Board of Appeals concluded that Refractory's patent fell short of the enablement requirement of Art. 83 EPC and had to be revoked. This case created quite a reaction, with many blog posts arguing that a number of details that AI patents need to disclose can be deduced from this decision. Specifically, it was argued that to be sufficiently disclosed, AI-related cases need to disclose at least the following aspects of the underlying model:
The specific machine-learning model or architecture used;
Precisely defined input and output variables;
The nature, quantity and quality of the training data and specifics on the training;
Guidance on how the invention can be carried out across the full breadth of the claim (which may mean providing multiple enabling embodiments if multiple models or data types are included).
In a recent, largely overlooked decision (T 0048/24, Ebara), the EPO clarified, that these criteria are not universally applicable, but applied only to the circumstances of that particular case, T1669/21:
"2.5 Similarly, in T 1669/21, the competent board indeed considered what the appellant referred to as "four criteria", namely that the patent did not disclose any specific example for carrying out the invention, which specific machine learning model and combination of specific input values could be used for predicting the claimed output parameter, and how the invention could be carried out over the whole claimed breadth. However, in the present Board's view, these aspects, in that particular case, illustrate the glaring gap between the breadth of the claimed invention and the level of detail in the patent, and do not constitute generally applicable criteria required for sufficiently disclosing a machine learning invention."
2.5 Similarly, in T 1669/21, the competent board indeed considered what the appellant referred to as "four criteria", namely that the patent did not disclose any specific example for carrying out the invention, which specific machine learning model and combination of specific input values could be used for predicting the claimed output parameter, and how the invention could be carried out over the whole claimed breadth. However, in the present Board's view, these aspects, in that particular case, illustrate the glaring gap between the breadth of the claimed invention and the level of detail in the patent, and do not constitute generally applicable criteria required for sufficiently disclosing a machine learning invention.
In other words, the above criteria are not universal, but the decision regarding sufficiency of disclosure needs to be made on a case-by-case basis.
That's good news!
Specifically, the Board notes that there is no requirement to provide a "specific example" of the claimed invention, i.e. input/output data, "details on the implementation and training of an exemplary machine learning model, or any information on the achieved accuracy of estimation". (The Board does note, however, that providing such an example might have been helpful.)
The bad news is that Ebara's patent was still revoked, and the argumentation by which the Board reached its conclusion. That is, the Board concluded that the invention did not fulfill the requirement of sufficiency of disclosure (Art. 83 EPC) across the whole breadth of the claim.
But let's look at the invention first. The invention deals with a device that uses a trained machine-learning model to estimate the composition of waste stored in a waste pit from data of a captured image of that waste.
Simplified somewhat, the invention claims:
A device comprising:
a unit that generates training data associated with data of an image of waste stored in a waste pit,
a unit that trains a model using that training data, and
a unit that inputs data of a newly captured image into the trained model to obtain a value representing composition of the waste shown in that image.
The idea behind the invention is that different types of waste (e.g. wet vs. dry; lots of paper vs. lots of plastic vs. lots of textiles) lead to different caloric values that can be obtained by burning the waste, and that this information can be extracted from images of the waste using a trained model.
Two expressions that the Board discussed in greater detail were "data of [an] image" and "composition of waste". Regarding the first term, the Board noted that the description uses only the term "image data" and argued that this expression was deliberately broadened in the claims. The Board interpreted this term as not being limited to a raw captured image, but as encompassing any data derived from it, including extracted features or even a simple averaged property such as color tone.
Also regarding the other expression, "composition of the waste", the Board adopted a broad reading and held that it is not limited to any specific value or degree of granularity, but covers everything from a coarse, human-comparable three-label classification of calorific value up to a precise numerical index resolved to 1%, since neither the claim nor the description ties the term to any particular accuracy requirement. (As a sidenote, it is noted that in the original Japanese expression, 廃棄物の質, might have been better translated as "quality of the waste".)
Thus, the Board concluded that, given the breadth of both the possible inputs ("data of an image") and the possible outputs ("composition of waste"), combined with the complete absence of guidance on which machine-learning model or architecture should be used to connect the two, the patent left the skilled person with an unmanageable number of parameter combinations to explore without any indication of which would actually work, amounting to a research program rather than routine implementation, and therefore insufficient across the whole breadth of the claim.
Now, sufficiency "across the whole breadth of the claim" is a standard that has traditionally been associated with the so-called "unpredictable arts" (in particular chemistry and biology) where small variations in a compound or biological system can produce unpredictable results, and where the EPO has long required applicants to make it plausible that the claimed effect is achieved across the entire scope claimed. Computer-implemented inventions, by contrast, have generally been treated as belonging to the "predictable arts", where a single worked example is usually enough to support a claim of corresponding scope. The fact that this "whole breadth" scrutiny is now being invoked so forcefully against AI inventions suggests that AI inventions may increasingly be treated by the EPO as sitting closer to the unpredictable arts than to conventional software. If that trend continues, patent attorneys drafting and prosecuting AI inventions may need to adopt the same drafting discipline (multiple worked examples, fallback positions, and narrower claim scope) that has long been standard practice in chemistry and biotech. This would be an unwelcome development for the field.