Why AI Seems Omnipotent — and Where Its Capability Boundaries Reappear in Real Work

A capability-boundary observation based on a failed AI-driven business experiment*

Scope note: This was not a controlled experiment, and it cannot represent every model, user, or industry. It is a first-hand account of a real project that failed, and of what that failure taught me about the difference between participating in a workflow and being able to own it.

I recently ended an attempt to build an AI-driven original Lolita fashion business.

The business logic was simple: design dresses, manufacture them, sell them, and make a profit. My plan was to use GPT and Lovart AI to take over much of the fashion designer’s role, shorten the patternmaker’s schedule, and assist with costing, fabric sourcing, and production preparation.

This expectation may sound aggressive, but it was not invented out of thin air. GPT could analyze existing designs, organize market data, write design narratives and briefs, and even produce DXF files. Lovart could generate fashion images. AI appeared able to enter almost every part of the project.

The project still failed.

It did not fail because AI could do nothing. The more interesting problem was almost the opposite:

AI could participate in nearly everything, but it could not reliably own any of the core tasks on which the business depended.

That experience led me to a broader question. If model capability boundaries are real, why does AI still appear omnipotent? And why do companies increasingly expect it to let fewer employees produce more than before?

1. A business that appeared well suited to AI

My original expectation: replace much of the designer’s role

In the design stage, I wanted AI to perform work traditionally handled by a fashion designer: selecting themes, developing individual garments, writing their narratives, and perhaps building a complete collection. I would set the commercial direction, provide references, evaluate the results, and make the final decisions.

This did not initially appear far outside the advertised scope of general-purpose and visual models. Theme development, storytelling, research, and design communication are language and image tasks. GPT can process large amounts of text and visual material, while Lovart can generate fashion concepts and illustrations.

I did not expect the first generation to be production-ready. My minimum expectation was more modest: after I supplied domain knowledge, calibrated the model’s aesthetic judgment, and ran several rounds of revision, it should produce at least one original Lolita design that met basic commercial requirements and justified continuing the project.

At first, GPT performed well.

I used it to collect publicly available fashion designs and analyze why they worked: their structures, proportions, visual language, and overall coordination. I also used GPT to build a sales-data extractor that organized public product information such as prices, sales signals, and other market indicators.

I am not claiming that its aesthetic analysis was scientifically validated, or that public sales data perfectly represented consumer preference. But at the task level, GPT did what I asked and did it well.

The failure appeared in the next step: turning analysis of existing work into an original design of my own.

The same model that could explain why an existing design looked good could not reliably use those explanations to create another coherent design. Its performance was poor in three connected areas: theme selection, narrative development, and the resulting design.

At the theme stage, GPT proposed subjects such as funerals and widowhood. Those themes are not categorically impossible in Gothic fashion, but they were high-risk and poorly matched to my intended brand and customer. The model offered them without a convincing commercial or creative reason, suggesting that it had not formed a stable understanding of the project’s thematic boundaries.

Its narratives often had complete structure but weak internal logic. One brief, translated here, was titled The Blue Ribbon Before Going Out:

Primary theme: Beautiful dresses should not have to wait for a special occasion
Sub-series: An outfit for an unplanned outing
Product position: Mature Sweet, everyday high-waisted JSK
Garment proposition: There had been no plan to go out, but after tying the blue ribbon on the dress, she decided to leave the house.

The brief then specified the silhouette, piping, pockets, and color palette. The text was not meaningless, but “tying a ribbon and deciding to go out” was only a thin everyday moment. Why did that story require those particular structures? The elements were written after the narrative; they did not grow from it.

The professional-looking hierarchy and detailed design language made the concept appear mature, while the theme, narrative, and garment were not meaningfully connected.

I understood that a general model could not naturally possess deep knowledge of every industry, especially one with high demands on silhouette, ornament, and coordination. I therefore did not treat the early failures as decisive.

I supplied specialized references, organized domain-specific Skills, ran blind selections and grouped comparisons, provided positive and negative examples, and conducted manual side-by-side evaluations. I wanted it to recognize when a complete design worked, not merely learn a vocabulary of bows and lace.

The result remained disappointing. The model could organize knowledge, analyze existing garments, and output polished themes, stories, and briefs. It still could not reliably transform those components into an original design that met my minimum bar.

In one week, I used my entire weekly GPT Pro allowance. In the same week, I spent more than 8,500 Lovart credits on design generation. At the rate I was paying — roughly USD 9 per 1,000 credits — the Lovart cost alone was about USD 76.50.

That amount is not expensive compared with hiring a professional designer. The more important cost was my time: judgment, feedback, repeated calibration, regeneration, and comparison. After that investment, I did not require AI to match the best professional designer. I required one result good enough to justify continuing.

It did not produce one.

Three Lovart workflows, no usable design

Because this article is mainly about general-model capability boundaries, I will keep the Lovart portion brief. I tested three workflows:

  1. Codex organized the task and called Lovart through its API. This produced the worst results.
  2. GPT wrote the prompt, and I manually transferred it into Lovart Agent. This was somewhat better.
  3. I rebuilt the context inside Lovart Agent, supplied references, summarized the intended style and constraints, and created a Lolita-specific Skill there. This was the strongest of the three.

Even the third workflow repeatedly produced bizarre construction, duplicate ideas, and uncontrolled color. It required constant regeneration and correction, which explains the 8,500-credit consumption. None of the three workflows produced a design I could use in the business or sell as an independent original concept.

Connecting two models did not simply add their capabilities together: intent, context, and constraints could be lost in transfer. More importantly, even the best configuration did not cross the acceptable delivery threshold. This does not prove that AI can never design Lolita clothing. It shows that even after narrowing the task to one usable original design, I still did not obtain a first acceptable result.

My second expectation: shorten the patternmaking schedule

I then tested whether AI could help with sampling and production preparation.

Traditional garment sampling involves patternmaking, fabric purchasing, sample production, fitting, and revision. I hoped AI could shorten that stage, or eventually perform much of the patternmaker’s job.

GPT could already generate DXF files, so this expectation did not seem completely detached from its visible capabilities. It could produce a file readable by industrial software. I wanted to test whether it could also produce a pattern that a factory could execute.

In retrospect, that inference crossed an important gap: generating a professional file format is not the same as performing the profession behind it. But the purpose of the experiment was to test that gap in practice rather than assume the answer in advance.

BOM: I did not require accuracy, but I did require consistency

I did not expect an accurate early bill of materials. Cost depends on patterns, fabrics, trims, processes, order quantities, and sample revisions; a stable BOM normally appears only after sampling.

But an estimate should still obey basic logic.

Earlier in the project, GPT had analyzed public pricing and cost information and helped establish a target retail range. During the sampling stage, its estimated BOMs often approached or exceeded that retail price, without identifying the conflict with its earlier conclusions.

I could not tell whether the original market analysis was wrong, the later BOM was wrong, or the product was impossible at the target price. GPT did not stop to reconcile two conclusions that could not both be true.

I did not require it to know the real BOM in advance. I required it to notice that its own estimates contradicted each other.

This was not primarily an offline-supply-chain problem. It was a project-consistency problem: a model can answer each question in a workflow separately without preserving the commercial relationships among those answers.

DXF: it generated an “industrial pattern” file, but not an industrial pattern

The DXF result initially surprised me in a positive way. GPT did not merely discuss patternmaking. It produced DXF files that could enter factory software.

I approached three factories. All three said that the files could not be used directly, and their feedback contained two independent findings.

First, the pieces and manufacturing information were not organized clearly enough for the factory to understand. Seam lines, allowances, notches, grain lines, numbering, piece relationships, and construction information did not form an executable industrial handoff.

Second, even if the communication problems were ignored, the pattern itself would not produce the intended Lolita dress.

The first failure might be improved with more standards, examples, and professional feedback. The second was more fundamental: the model had not translated the target silhouette, proportions, volume, and structural relationships into the pattern.

From GPT’s perspective, the task was complete because it had generated and delivered a DXF. From the factories’ perspective, the task was incomplete because the file was neither usable nor structurally the intended garment.

AI successfully generated a file that looked like the output of industrial patternmaking. It did not successfully generate an industrial pattern.

In theory, GPT could generate a draft for a patternmaker to repair. But that would not remove the patternmaker or necessarily shorten the schedule; it might instead add an unfamiliar artifact to interpret. A real AI patternmaking pipeline would require qualified pattern data, industrial standards, professional feedback, and specialized software — beyond my expertise, capital, and project scope.

GPT crossed the boundary of generating a DXF file. In my project, it did not cross the boundary of the patternmaker’s profession.

Fabric: the real-world gap was much larger than expected

GPT did warn me that it had no direct connection to offline fabric inventory and that generated garment images might differ from real materials. I understood and accepted that limitation.

Based on its assessment of the images, however, the design still appeared achievable with commercially available bulk fabric. I expected differences in hand feel, color, sheen, and drape, or that I would need to choose among several suppliers offering the same general material.

I took the AI-generated garment images to offline fabric sellers. Most told me that they did not carry materials capable of producing that result. To preserve the appearance shown in the images, I would need custom-developed fabric.

That was not an ordinary price or quality difference. It changed the procurement route from buying bulk material to custom development, introducing fees, sampling, minimum orders, longer lead times, and new risks. It also destabilized the BOM, retail price, capital requirement, and schedule.

I cannot prove that the fabric was unavailable everywhere. My purchasing channel, location, order size, or communication may have contributed. I searched as far as I reasonably could.

For a business, “it may exist somewhere” is not equivalent to “I can buy it within my cost, lead-time, and minimum-order constraints.” GPT identified an information risk, but did not convey how far that risk might move the plan. I expected to choose among materials of the same type; I encountered a move from bulk purchasing to custom manufacturing.

Moving from software into physical production

I did not conduct a full production-stage experiment, so this is factory feedback rather than a verified conclusion. The factory said direct AI involvement in cutting, sewing, and pressing might require dedicated hardware because existing workflows were designed around human operators.

AI could assist with scheduling, records, quality information, or production management. But replacing physical labor would require machine vision, sensors, actuators, and compatible production equipment — not simply one more Agent connected to an office computer.

This exceeded both my capabilities and the original capital and organizational design of the business. At that point, the AI-driven original Lolita design and production project ended.

2. What this case can — and cannot — establish

AI contributed real work: it analyzed public designs, organized sales data, created briefs, estimated BOMs, generated DXF files, and produced realistic-looking garment images. Judged by local outputs, it entered nearly the entire value chain.

It nevertheless failed at the points that determined whether the business could proceed: original designs did not reach the minimum bar after extensive calibration; cost conclusions were inconsistent across stages; patterns could not be used by factories and did not produce the target garment; images were disconnected from an accessible procurement path; and deeper entry into physical production might require a new hardware system.

These failures were not all equally strong evidence. An early BOM is inherently uncertain. GPT had warned about fabric risk. The production-hardware point came from a factory rather than my own full test. The clearest failures were original design and patternmaking.

This case cannot prove that GPT is unsuitable for fashion, exclude my own professional limitations, or predict what future AI will do. With designers and patternmakers retained, AI may still create real value through research, discussion, documentation, and early drafts.

What the case does establish is narrower, but sufficient for a commercial decision:

Under the quality requirements, resources, and tools available in my project, AI could not own the core responsibilities of original design and sampling.

A company does not need to prove that a technology will be impossible forever before it is allowed to stop investing. If the first acceptable result does not appear within the time, money, labor, and risk the company can afford, it can reasonably conclude that the path is not currently suitable for that project.

AI participated across the workflow, but it did not complete the workflow. I did not abandon the project because of one bad generation. I abandoned it because the capability boundary reappeared every time the work moved from “this looks promising” to “this must now work in reality.”

3. Why companies expect AI to be omnipotent

After the experiment failed, a more difficult question remained: why had I believed that a general-purpose model could let one person perform work previously distributed across several professions?

The expectation comes partly from what general models genuinely do, and partly from corporate demand for cost reduction and efficiency.

In a low-growth and uncertain environment, consumers and businesses become more cautious. In mainland China, retail sales grew by 1.2% year over year in the first seven months of 2026, private investment fell by 9.4%, and real-estate sales by value fell by 13.1%. The National Bureau of Statistics also noted weak demand and operational pressure on some businesses.

This does not mean every company’s profit was falling: industrial profits still grew, with large differences among sectors. But revenue growth has become harder for many firms dependent on domestic demand or intense price competition.

The basic relationship is simple:

Profit = Revenue − Cost

When revenue cannot grow quickly, cost management becomes the most immediate lever available.

Labor is not the largest cost in every industry or the only adjustable one. But it exists across departments, is easy to quantify, and can affect financial results quickly. When production and supply-chain improvements require more time, headcount becomes one of the easiest variables to change immediately.

“Cost reduction and efficiency improvement” is not new, but became especially visible in China’s technology sector around 2022. Alibaba’s improvement in adjusted EBITA through cost and efficiency measures gave later expectations a real reference point.

Traditional layoffs and Agent-era efficiency programs may use the same language, but they contain different ambitions.

Traditional layoffs remove or redistribute work to maintain roughly the same output with a smaller payroll. With Agents, remaining employees are expected to use models to perform tasks previously distributed across several roles.

The objective changes from:

Use fewer people to maintain existing output

to:

Use fewer people plus AI to produce more than before.

That expectation is not purely imaginary. Microsoft’s 2025 Work Trend Index reported that 33% of leaders were considering headcount reductions, 78% were considering hiring for new AI roles, and 83% believed AI would allow employees to take on more complex and strategic work earlier. These expectations appearing together illustrate the combined ambition: reduce conventional labor, increase AI capability, and raise output per employee.

This creates a competitive ratchet.

If one company uses AI to reduce unit cost, shorten delivery time, or multiply content and product output, competitors no longer face a neutral technology choice. The logic becomes:

A competitor increases output with AI
→ market expectations for speed, volume, and price rise
→ refusing to follow may reduce competitiveness
→ investment becomes necessary even before the return is clear

AI adoption may not reduce total cost. Companies must buy tools and compute, build systems, train staff, and expand review and security. Once one task becomes cheaper, competition may demand more tasks, content, and service.

Companies can become more efficient without becoming more profitable. Gains may flow to customers through lower prices and faster delivery, or to model and compute providers. Yet the risk of not adopting remains.

A company may adopt AI not because AI has already proved that it will create more profit, but because the company cannot afford the possibility that a competitor has proved it first.

From local success to unlimited expectation

Two additional forces strengthen the expectation.

The first is the word general. These systems work with text, code, images, data, research, communication, and creative tasks. When an “AI assistant” appears able to enter most knowledge-work domains, “it can help with many jobs” easily becomes “it can replace many roles.”

The second is that AI has produced real gains. In OpenAI’s enterprise survey, 75% of surveyed workers said AI improved speed or quality across several departments. The provider and customer sample cannot represent the whole economy or prove headcount reduction, but local productivity gains are real. (OpenAI, The State of Enterprise AI)

But expectations rarely stop after the first gain.

If AI reduces the time required for one task, can it reduce headcount? If it can reshape one role, can it reshape a department? If it works in internet companies, can it enter traditional industries? If one small team completes work that previously required a large team, can every company reproduce that leverage?

Each local success expands the scope of what must be proved next.

For the “smaller cost, larger return” thesis to work, at least two conditions must hold. Employees must develop broad AI operating capability — not merely prompt writing, but model selection, task decomposition, workflow design, context maintenance, evaluation, error detection, and business integration. And the models themselves must cover a sufficiently large portion of the work, or companies must believe that employees can keep pushing the boundary outward with Skills, tools, and workflows.

This brings us to the central question:

Do models really have no capability boundary?

4. Participating in every task is not the same as owning every task

Models obviously have capability boundaries. The difficulty is that these boundaries are not clear and stable like those of conventional tools.

They resemble an uneven, moving region. Change the task and the boundary moves. Change the wording and the result changes. Add tools and it expands. Remove context and it contracts. Raise the standard from “succeeded once” to “can deliver reliably,” and it contracts again, often dramatically.

The central illusion of general-purpose models is this:

They can participate in almost every task, but they cannot reliably own every task.

We need to distinguish at least three levels:

responding to a task
≠ producing a plausible-looking result
≠ reliably and independently delivering a result a business can accept

I previously worked at a leading game company serving global markets, with AI scale advantages and dedicated production workflows that an individual entrepreneur could not reproduce.

In one video workflow, the team used Claude and Codex to translate a director’s storyboard or an already-defined commercial requirement into prompts that Seedance could interpret more effectively. When generation failed, humans described the error to Claude or Codex, which rewrote the prompt before another round of generation and review.

One frontier model was effectively translating into the “model dialect” of another frontier model. The final output came from a system:

business or creative staff

  • Claude / Codex
  • Seedance
  • human review
  • repeated rework

This workflow may be rational and productive. But it creates an attribution problem. The fact that the system completed a video does not mean Seedance completed it independently. Claude or Codex helping it cross part of a boundary does not mean the original boundary did not exist.

I was therefore not rejecting AI after my first bad prompt. I had seen it operate with scale, specialized processes, cross-model translation, human review, and repeated correction.

I do not claim that my personal method was optimal. Some failures may still reflect my limitations. But even with that experience, the communication, calibration, translation, verification, and rework required to complete a concrete business task far exceeded my expectation. In my own career, the intensity of this friction was unprecedented.

“The user does not know how to use AI” may explain part of the failure. It cannot fully explain why a user familiar with scaled AI workflows, willing to build Skills and test multiple model configurations, still faced such a high adaptation cost.

If humans supply context, decompose tasks, repair outputs, translate between models, and accept the final risk — while every successful result is attributed simply to “AI” — the model will appear stronger than it is on its own.

5. Why capability boundaries disappear from view

Breadth hides depth

GPT could discuss aesthetics, sales data, stories, BOMs, patterns, fabrics, and production. It was not ignorant, and it produced real value in several places.

But participating across many stages and reaching a professional delivery threshold at every stage are different capabilities. Anthropic’s research on actual AI use similarly suggests that AI participates in parts of many occupations far more often than it covers most tasks within an occupation. Today’s reality is closer to “touching many jobs” than “fully taking over many jobs.”

Professional form hides professional substance

A well-structured design brief does not mean the story and design work. A DXF that opens does not mean a factory can use it. A realistic garment image does not mean its material exists within the actual supply chain.

Models are very good at producing the visible form of professional output. Businesses need the professional relationships behind that form to hold in reality.

System performance is attributed to the model

When a result is produced by a user, several models, Skills, search, APIs, human review, and offline professionals, it is easily described as “AI did it.”

The model may produce the visible 80%, while employees provide the remaining 20% through context, repair, rework, and risk ownership. Yet the most expensive professional judgment may be concentrated in that final 20%.

If the human labor used to fill the boundary is omitted from the cost calculation, the entire result is credited to AI.

Failure is reclassified as insufficient AI skill

When a model fails, companies can ask users to write better prompts, build more Skills, switch models, redesign the workflow, or add another model for translation and verification.

These approaches sometimes work. But if every failure is first interpreted as “the user does not know how to use AI,” the model’s boundary can never become visible. The boundary has merely been transferred to the user.

Humans can keep moving the target

If AI cannot produce an original design with sales value, the goal can be changed to “design inspiration.” If it cannot independently make patterns, the goal can become “provide a draft to the patternmaker.” If it cannot provide a reliable BOM, it can provide an “early estimate.” If it cannot enter physical production, it can participate only in production management.

All of these may be sensible business decisions. They prove that AI is useful within narrower tasks. They do not prove that the original capability boundary disappeared.

6. Conclusion: a system completing the task does not mean the model has no boundary

This article is not intended as an attack on GPT, nor as proof that AI has no value in fashion.

The project was an attempt to apply the “cost reduction, efficiency improvement, and asymmetric leverage” thesis to a small business: use general models and AI design tools to replace or shorten original design, patternmaking, and sampling preparation, and create a Lolita product line with less labor and capital.

The result was clear. Under the requirements and tools available to me, it failed.

I cannot prove that my AI use was theoretically optimal. But I had worked in a leading game company with scaled AI workflows, and in the personal project I supplied extensive domain knowledge, created Skills, conducted blind and grouped comparisons, and tested several combinations of GPT, Codex, Lovart Agent, and API calls.

After that level of effort, “the user just needs to write better prompts” is no longer an adequate explanation of the outcome.

I could add more Skills, workflows, models, and human review, lower my standards, or redefine the goal from “AI owns design” to “AI helps organize references and discussion.” Those changes might allow a different project to continue.

But none would prove that the model had no boundary. The first approach means that humans contribute more labor to fill the gap. The second means that I have changed the task the model was originally expected to perform.

AI appears omnipotent because it can enter almost any task and always generate another next step. When it cannot own the outcome, people can add prompts, Skills, tools, models, and human review — or lower the objective — until the larger system produces something.

But:

A system eventually completing a task does not mean the model has no capability boundary. The human compensation required to make the system work should not be renamed as model capability.

If proving that AI is omnipotent requires humans to keep learning how to accommodate it, or to keep redefining what they originally wanted to accomplish, then what has been demonstrated is not that AI has no boundaries. It is that humans can keep moving the target.


Discussion question: In your own work, where have you seen the largest gap between an AI system being able to participate in a task and being able to own the result reliably? How do you account for the human labor used to close that gap?

There are probably many reasons, but the first that comes to mind is the echo-chamber effect.

When people repeatedly encounter impressive examples of what AI can do—especially in communities where those successes are frequently shared and amplified—it can become easy to overlook the failures, limitations, and cases that never get posted.

That can create an impression that AI is far more capable, reliable, or even “omnipotent” than it actually is.