🔍 Read the full analysis: Could You Create The AI Model You Can't Find? on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A Hugging Face contributor reports using the ML-intern agent to build and publish seven custom models over several days, including a small prompt rewriter and a citrus disease classifier. The contributor reported low compute costs and gains on selected tests, but the results are self-reported and do not establish how the workflow performs for other users or tasks.
A Hugging Face contributor says the platform’s ML-intern agent helped build and publish seven custom models over several days, as described in the original analysis, including a 0.8-billion-parameter prompt rewriter and a model for identifying citrus problems in images. The examples offer a practical account of agent-assisted model development, but the performance figures and compute costs are the contributor’s own reports, not an independent evaluation.
The work began with a request for a smaller version of the prompt rewriter included with Qwen-Image 2.1. The contributor said the original model has 9 billion parameters, needs about 20 GB of memory and produces thousands of tokens before returning a paragraph. After finding compressed versions of the large model but no smaller alternative, the contributor used the agent to create a 0.8B model. They reported valid output 99.7% of the time and about one-quarter the token use of its teacher model. The reported compute bill, including teacher-generated labels for 8,797 requests, was about $16.
Another project fine-tuned Qwen3.5-2B to classify citrus pests, diseases and nutritional deficiencies from photographs. The contributor said the training set had 3,017 annotated images across 21 categories. On 335 test photos, they reported that accuracy for identifying the correct problem rose from 14.9% for the base model to 52.8% after fine-tuning. Training ran for two epochs on one A10G GPU, with reported compute costs of about $1.90.
The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the agent generated 24,722 transparent images of scanned household objects across 24 angles. The contributor said training took about 90 minutes on one A100 and the project cost about $16 in compute, including failed jobs that had to be resubmitted. The account says model cards and evaluations were published on Hugging Face, while details for all seven projects were not included in the supplied description.
A Lower Barrier to Custom Models
The examples suggest that an AI agent can take on parts of a workflow that otherwise require a developer to coordinate data preparation, test runs, training, evaluation and publication. For people with a narrow need—such as a smaller text tool or a classifier for a specific crop—the reported compute bills indicate that experimentation can sometimes be inexpensive. Those figures cover compute as described by the contributor, however, and do not provide a full accounting of data preparation, human review or time.
The citrus example also shows why a comparison matters: the contributor reported results for the fine-tuned model alongside the base model on the same 335-image test set. That is more informative than a standalone score, but it does not show whether the result holds on other photographs or in field conditions. The character LoRA account offers a separate caution: the contributor said later checkpoints began affecting prompts unrelated to the character. A model can therefore become less useful beyond its target task, even when training appears to improve the intended output.
As an affiliate, we earn on qualifying purchases.
From HuggingChat to Published Models
According to the contributor, each project began as a message in HuggingChat with ML-intern enabled. The agent proposed a plan, sought approval before paid work, ran a small test, and then carried out training, evaluation and publication using Hugging Face hardware. When a prompt did not specify a budget, the agent offered spending options and asked the user to choose.
The contributor said their instructions grew more detailed during the sequence, from about 450 words for the first project to nearly 2,000 by the sixth. They included the dataset, base model and training script, and asked for a baseline, a small test run and a spending cap. The contributor’s prompts are said to be available in a public GitHub repository. This is a description of one person’s process, rather than an independent assessment or a guarantee of comparable outcomes for other users.
“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”
— The Hugging Face contributor
As an affiliate, we earn on qualifying purchases.
How Reliable Are the Reported Gains?
The supplied account does not include independent replication, complete evaluation protocols for every model, or detailed descriptions of all seven projects. It is not clear how the 99.7% valid-output rate was measured, how test examples were selected, or whether the citrus test images were independently reviewed. Results on the reported test sets also do not establish performance on new data, other hardware or different budgets.
The cost figures are presented as compute costs, not a full project budget. They do not establish the expense of preparing and checking datasets or the time required to write prompts and review outputs. The contributor’s account also does not show whether other users would receive similar plans or results from the agent. These limits mean the examples illustrate a possible workflow, but cannot establish that its reported costs or gains are typical.
Hugging Face model development tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
More Replication Needed
The contributor says the seven models, their evaluations and project prompts are available through Hugging Face and GitHub, giving readers material to inspect. The next useful test would be comparisons by other users across different datasets and tasks, with published evaluation methods and clear accounting of costs. Independent checks could help show whether the results extend beyond the contributor’s own projects.
For future work, the contributor’s stated process calls for checking a base-model score before training, running a small trial, confirming that saved weights changed, and setting a spending limit. Those are described safeguards, not proof that a project’s data or outputs are sound. Until broader evaluations appear, readers should treat the reported numbers as project-specific results, not expected performance for every ML-intern user.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is ML-intern in this account?
It is the agent the contributor says they used through HuggingChat to plan and carry out model-development tasks, including test runs, training, evaluation and publication.
What models did the contributor report making?
The account describes a 0.8B prompt rewriter, a citrus pest and disease classifier, a character-generation LoRA, and a camera-angle LoRA, among projects said to total seven models. The supplied details do not describe every project.
Are the reported performance numbers independently verified?
No independent verification is described in the supplied account. The reported scores and costs come from the contributor, and full evaluation methods are not provided for every model.
How much did the projects cost?
The contributor reported about $1.90 in compute for the citrus classifier and about $16 for the prompt rewriter and camera-angle project. These figures are not a complete accounting of labor, data preparation or all project expenses.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
