From Prompt Language to Commercial Illustration Workflow: A Computational Review and Prompt-Control Framework for Generative AI Design ()
1. Introduction
Commercial illustration is not merely a matter of producing attractive images. It is a situated design practice that translates client briefs, brand identities, audience expectations, cultural references, production constraints, and revision cycles into visual form. The diffusion of text-to-image models has altered this practice because language can now initiate visual production at scale. A designer can move from a short prompt to dozens of alternative images, then select, edit, combine, and re-prompt.
The literature on AI and creative production has developed along several partially separate lines. Broad reviews situate AI within the creative industries and emphasize opportunities and risks for cultural production (Anantrasirichai & Bull, 2021). Technical studies explain the language-image foundations of contemporary generative systems, including contrastive language-image pretraining and latent diffusion models (Radford et al., 2021; Rombach et al., 2022). Dataset studies such as DiffusionDB and public prompt collections make it possible to examine prompt culture empirically (Gustavosta, 2022; Wang et al., 2022).
For commercial illustration, the most important issue is not whether AI can generate an image. The deeper issue is how AI reorganizes the design process. A prompt is not a simple text description. It encodes subject matter, medium, detail level, composition, genre, atmosphere, platform aesthetics, and sometimes direct references to artists or styles. Prompt writing therefore becomes a new site of design judgement. At the same time, the designer must still evaluate whether the output fits a brief, communicates a message, satisfies a client, and avoids ethical or legal risks (Rezwana & Maher, 2023).
Existing studies have addressed co-creative drawing workflows (Lawton et al., 2023), designer agency (Guo et al., 2023), character design and illustration prototyping (Ling et al., 2024), design-space exploration (Davis et al., 2024; Suh et al., 2024), and visual design applications (Huang & Zheng, 2022). Education studies also show that prompt engineering and iterative AI use are increasingly important in art and design learning (Cotroneo & Hutson, 2023; Fathoni, 2023; Hutson & Lang, 2023; Vartiainen et al., 2023).
This paper addresses three research questions: 1) What conceptual and methodological roles do thirty core references play in explaining generative AI for visual design and commercial illustration? 2) What prompt-control dimensions are visible in illustration-related text-to-image prompt statistics? 3) How can literature evidence and prompt-language evidence be integrated into a workflow model for AI-assisted commercial illustration?
2. Literature Foundation
2.1. Prompt Engineering as Design Method
Prompt engineering is central to AI-assisted image creation because it transforms design intent into machine-interpretable language. Liu and Chilton (2022) show that prompt writing is a design activity involving iteration, specificity, visual descriptors, and model feedback. Sanchez (2023) further shows that text-to-image users form a community of practice in which prompts circulate as reusable design knowledge. Cotroneo and Hutson (2023) and Herath et al. (2026) extend this point by linking prompt engineering to creative learning and interface support.
2.2. Human-AI Co-Creation and Designer Agency
A second body of work examines the relationship between designers and generative systems. Collaborative diffusion frames generative AI as a means of designerly co-creation rather than simple automation (Verheijden & Funk, 2023). Lawton et al. (2023) show that co-creative drawing involves emergence and control, while McCormack et al. (2020) identify timing, feedback, and control as central interaction-design problems in creative AI collaboration. Guo et al. (2023), Vartiainen et al. (2024), Wang et al. (2025), and McGuire et al. (2024) jointly suggest that AI-assisted commercial illustration should be analysed as a distributed workflow in which control moves between human intention, prompt language, model output, and post-editing judgement.
2.3. Commercial Illustration, Design Education, and Data Foundations
Commercial illustration sits at the intersection of creative industries, advertising, cultural translation, visual design, and client-facing production. Hanna (2023) connects Midjourney to artistic and advertising creativity. Lyu et al. (2023) show how generative AI can mediate cultural translation in jewelry design, and Wang and Zhang (2023) examine adoption of generative AI for art design among Chinese Generation. Du et al. (2023) add that human-AI image-making can involve affective and reflective dimensions. The technical basis of text-to-image design relies on models and data infrastructures that connect language, images, prompts, affective descriptions, and scholarly metadata (Achlioptas et al., 2021; OpenAlex, 2026; Radford et al., 2021; Rombach et al., 2022; Wang et al., 2022).
Table 1 summarizes the theoretical constructs that organize the focused synthesis.
Table 1. Theoretical constructs and their operational roles in the study.
Figure 1 maps the thirty-reference evidence base by publication year, theme, evidence role, and methodological use in the synthesis.
Figure 1. Thirty-reference evidence atlas. Panel A maps the thirty references by publication year and thematic role. Panel B summarizes evidence roles across the corpus. Panel C shows how the evidence base supports the research questions and the final framework.
3. Methods
3.1. Focused Review Source Selection
The thirty core references were selected through a focused synthesis rather than a claim to exhaustive systematic review. Search sources included Google Scholar, ACM Digital Library, IEEE Xplore, ScienceDirect, SpringerLink, Taylor & Francis Online, Hugging Face dataset pages, and OpenAlex metadata. Search strings combined terms such as generative AI, text-to-image, prompt engineering, Stable Diffusion, Midjourney, commercial illustration, visual communication design, graphic design, human-AI co-creation, design-space exploration, art education, and AI ethics. Inclusion required at least one direct contribution to a) prompt engineering or prompt culture, b) text-to-image technical or dataset foundations, c) human-AI co-creation and designer agency, d) commercial or visual design practice, e) art and design education, or f) ethical and disclosure issues in creative AI. General AI papers without a visual-design or image-generation connection were excluded. The synthesis uses categorical theme and evidence-role coding only; no numerical weighting scheme is used to calculate or rank the results.
3.2. Prompt Sample Construction
The prompt sample was built from the public Hugging Face dataset Gustavosta/Stable-Diffusion-Prompts, using the default configuration and train split. The access point was the Hugging Face dataset page and rows API endpoint for the repository Gustavosta/Stable-Diffusion-Prompts. No separate named version tag was available in the project files, so the dataset is identified by repository name, default configuration, train split, and access point. Rows were scanned sequentially from the train split and retained when the normalized prompt string contained at least one illustration-related keyword or phrase. Empty prompts, non-matching prompts, and exact duplicate normalized prompt strings were excluded. Whitespace was normalized before deduplication, and the first 3000 unique matching prompts were retained for aggregate analysis. The prompt text itself is not redistributed in the article; only aggregate statistics and scripts are included in the project folder.
The full sampling filter was: illustration, illustrator, poster, advertising, advertisement, brand, branding, logo, mascot, book cover, children book, children’s book, editorial illustration, commercial, packaging, package design, character design, vector, flat illustration, line art, concept art, digital painting, and artstation. This full filter is now stated to make the construction of the 3000-prompt sample reproducible.
3.3. Operational Coding of Prompt-Control Dimensions
Each prompt could be coded into multiple prompt-control dimensions because commercial illustration prompts often combine genre, style, quality, platform, and reference cues in the same sentence. Coding was binary at the prompt level: a prompt received 1 for a dimension if any indicative term or phrase for that dimension appeared in the lowercased prompt string, and 0 otherwise. Therefore, all reported percentages are calculated per prompt rather than per term occurrence.
Table 2 defines the five prompt-control dimensions operationally.
Table 2. Operational coding rules for prompt-control dimensions.
Dimension |
Indicative terms or phrases |
Coding rule |
Workflow interpretation |
Illustration core |
illustration, illustrator, concept, character, drawing, line, vector, concept art, line art |
Coded 1 if prompt specified an illustration genre, visual object, or drawing mode. |
Locates AI use in ideation, character exploration, storytelling, and prototyping. |
Commercial visual framing |
poster, brand, branding, logo, advertising, advertisement, book cover, packaging, package design, commercial |
Coded 1 if prompt explicitly named a market-facing visual communication task. |
Tests whether prompt language makes client-facing goals explicit. |
Style-quality intensification |
detailed, intricate, realistic, sharp, focus, cinematic, smooth, elegant, ultra |
Coded 1 if prompt used quality, detail, finish, or polish descriptors. |
Explains how users attempt to control visual finish before
post-editing. |
Platform aesthetics |
artstation, trending, digital painting, digital, painting, unreal, octane, render |
Coded 1 if prompt anchored output in online platform or digital-art conventions. |
Identifies possible aesthetic convergence around platform visual cultures. |
Human-reference styling |
named artists, Mucha, Artgerm, Greg Rutkowski, style of, by, historical style cues |
Coded 1 if prompt used human artist references or direct style-attribution cues. |
Triggers authorship, imitation, disclosure, and ethical-review concerns. |
3.4. Text Preprocessing and Co-Occurrence Modelling
Before frequency and co-occurrence analysis, prompt text was lowercased and whitespace-normalized. Multi-word phrases such as book cover, children book, character design, package design, concept art, digital painting, flat illustration, and line art were handled through phrase matching on the raw lowercased prompt string before tokenization. Token-level counts then used a regular-expression tokenizer that retained alphabetic tokens of three or more characters and allowed internal hyphens. A small stopword list removed function words and generic descriptors that did not serve the five analytical dimensions.
The co-occurrence unit was one prompt. For the network analysis, each selected network term was treated as binary within a prompt; an edge count increased by one when two network terms appeared in the same prompt. The adjacency matrix therefore contains raw prompt-level co-occurrence counts. The visualization displayed edges above a descriptive threshold of 13% of the maximum observed edge weight to reduce clutter. The analysis was implemented in Python using regular expressions and pandas for counting and matrix construction, with Matplotlib used for figure generation. No inferential statistical test was applied because the aim was descriptive mapping of prompt-control vocabulary rather than population estimation.
Table 3 summarizes the procedural details used for prompt sampling, preprocessing, and co-occurrence modelling.
Table 3. Prompt sampling, preprocessing, and co-occurrence procedure.
Procedure component |
Revised specification |
Dataset access point |
Hugging Face repository Gustavosta/Stable-Diffusion-Prompts; default configuration; train split; rows API access point. |
Sample construction |
Sequential scan of train rows; retain prompts matching the full keyword filter; normalize whitespace; remove exact duplicate normalized prompt strings; retain first 3000 unique matching prompts. |
Exclusions |
Empty prompts, prompts without any filter keyword or phrase, and exact duplicate normalized prompt strings. |
Multi-word phrases |
Phrase matching on raw lowercased text before tokenization, including book cover, character design, concept art, digital painting, line art, and package design. |
Tokenization |
Lowercase text; use alphabetic regular-expression tokens of three or more characters; remove a small stopword list. |
Percentage calculation |
Binary prompt-level presence; percentages report share of prompts containing at least one term in the dimension. |
Co-occurrence
modelling |
One prompt as unit; binary term presence; raw edge counts; visualization threshold at 13% of maximum edge weight; Python/pandas/Matplotlib implementation. |
3.5. Workflow Framework Construction
The final framework was constructed through abductive synthesis: the prompt evidence was interpreted through literature on human-AI co-creation, design-space exploration, and design education (Davis et al., 2024; McCormack et al., 2020; Suh et al., 2024; Verheijden & Funk, 2023; Wang et al., 2025). The aim was not to infer professional behaviour directly from public prompt data. Instead, the aim was to identify workflow implications that can guide future interviews, design experiments, and commercial case studies.
4. Results
4.1. Reference Base Organised around Method and Evidence Role
The thirty references form a layered evidence base rather than a flat list. Technical and data references cluster around 2021-2022, reflecting the emergence of CLIP, latent diffusion, and large-scale prompt datasets (Radford et al., 2021; Rombach et al., 2022; Wang et al., 2022). Prompt engineering, co-creation, and design application references expand strongly from 2023 onward, corresponding to the practical diffusion of text-to-image tools (Ling et al., 2024; Sanchez, 2023; Verheijden & Funk, 2023).
4.2. Prompt Language Is Dominated by Visual Control
Figure 2 visualizes the prompt-control landscape, including dimension prevalence, co-occurrence links, and term concentration by design function.
Figure 2. Prompt-control landscape. Panel A shows the proportion of prompts containing each control dimension. Panel B visualizes co-occurrence among major prompt controls. Panel C shows term concentration by design function.
The prompt aggregate statistics show that illustration-related prompt language is heavily organized around visual control. Illustration-core terms appear in 95.23% of the sample. Style-quality terms appear in 84.93%, platform-aesthetic terms in 84.80%, and human-reference terms in 77.70%. By contrast, explicitly commercial visual terms appear in only 8.17% of prompts. Table 4 reports these prompt-level percentages.
Table 4. Observed prompt-control dimensions in the 3000-prompt aggregate sample.
Dimension |
Share of prompts |
Design reading |
Illustration core |
95.23% |
Genre and object specification dominate prompt construction. |
Style-quality intensification |
84.93% |
Users heavily encode finish, polish, and visual detail. |
Platform aesthetics |
84.80% |
Online visual cultures structure the expected image look. |
Human-reference styling |
77.70% |
Style is frequently controlled through human artistic references. |
Commercial visual framing |
8.17% |
Commercial goals are present but linguistically under-specified. |
4.3. Alternative Explanation for Low Commercial Visual Framing
The low share of explicit commercial visual framing should be interpreted cautiously. One explanation is that many public prompts pursue visual polish, portfolio aesthetics, or fan-art conventions rather than client-facing communication. However, an alternative explanation is that commercial intent may be implicit in non-commercial wording. A prompt that says concept art, detailed digital painting, or cinematic poster-like composition may still be used for advertising ideation, brand mood boards, or product visualization even if it does not contain the terms brand, packaging, advertising, or logo. The public dataset may also bias the estimate downward because it captures open prompt culture rather than confidential client briefs or professional studio logs.
This balanced interpretation means that the finding should not be read as evidence that generative AI lacks commercial relevance. It indicates that explicit commercial language is sparse in the public prompt sample and that professional commercial intent may require additional data sources such as designer interviews, client briefs, or process logs.
4.4. Co-Occurrence Reveals a Platform-Style-Quality Loop
The co-occurrence network reveals a dense loop among concept, detailed, artstation, digital, painting, sharp, and focus. This loop indicates a common prompt logic: users specify a general genre such as concept art or illustration, intensify expected image quality, and anchor the output in platform aesthetics. In commercial illustration, this creates both affordances and risks. The affordance is rapid production of polished alternatives. The risk is stylistic homogenization: outputs may converge on widely circulated digital-art aesthetics rather than on brand-specific communication.
5. Discussion
5.1. From Image Generation to Workflow Redistribution
The results support a shift from thinking about generative AI as image generation to thinking about it as workflow redistribution. In a traditional illustration workflow, the designer’s labour is concentrated in sketching, composition, refinement, and manual execution. In an AI-assisted workflow, the designer’s labour is redistributed across brief interpretation, prompt translation, output curation, post-editing, and ethical review. This interpretation is consistent with research on co-creation and designer agency (Guo et al., 2023; Vartiainen et al., 2024; Wang et al., 2025).
The redistribution does not reduce the importance of design expertise. On the contrary, it makes expertise more visible at decision points. The designer must know how to write a prompt, but also when a prompt has failed. The designer must compare outputs, recognize visual cliches, maintain brand consistency, and decide whether a generated image is acceptable for commercial use. This supports McGuire et al.’s (2024) emphasis on self-efficacy in creative collaboration with AI.
5.2. A Commercial Illustration Workflow Model
Figure 3 proposes the final workflow model. The model contains six stages: brief interpretation, prompt translation, generative exploration, curation and comparison, post-editing and delivery, and ethical/client review. The model is intentionally not a linear automation pipeline. It includes feedback loops because commercial illustration is shaped by revision, client response, and visual evaluation. As summarized in Table 5, the proposed workflow translates insights from prior literature and prompt-based evidence into six interrelated stages of commercial illustration practice. The framework emphasizes progression from brief interpretation and prompt translation to generative exploration, human-led curation, post-editing, and ethical/client review. Overall, Table 5 highlights that generative AI functions most effectively as a supportive design tool embedded within a human-centered workflow, rather than as an autonomous replacement for professional creative judgment.
![]()
Figure 3. Vector framework for AI-assisted commercial illustration. The framework links reference coding, prompt aggregate analysis, co-occurrence logic, and ethical constraints to a six-stage commercial illustration workflow.
Table 5. Workflow-level theoretical inference from literature and prompt evidence.
Workflow stage |
Evidence basis |
Commercial implication |
Brief interpretation |
Creative-industry and advertising studies stress communication goals and client contexts (Anantrasirichai & Bull, 2021; Hanna, 2023). |
Start from brief, audience, brand tone, format, and revision constraints rather than model novelty. |
Prompt translation |
Prompt engineering studies define prompt writing as iterative visual specification (Herath et al., 2026; Liu & Chilton, 2022; Sanchez, 2023). |
Treat prompt versions, style terms, and image-ratio settings as design process evidence. |
Generative exploration |
Co-creation and design-space tools show the value of structured alternatives (Davis et al., 2024; Suh et al., 2024; Verheijden & Funk, 2023). |
Use AI mainly as an early-stage exploration engine, not as a final author. |
Curation and comparison |
Designer agency and experience shape evaluation under AI co-creation (Guo et al., 2023; McGuire et al., 2024; Wang et al., 2025). |
Keep human judgement central to originality, brand fit, relevance, and visual coherence. |
Post-editing and delivery |
Illustration prototyping and visual design studies connect AI output with production work (Huang & Zheng, 2022; Ling et al., 2024). |
Require craft correction, compositing, typography control, and deliverable discipline. |
Ethical/client review |
Ethical co-creativity literature highlights disclosure, imitation, and responsibility (Rezwana & Maher, 2023). |
Formalize style-reference checks, client approval, and copyright-risk review before delivery. |
5.3. Implications for Design Education and Commercial Practice
The findings have practical implications for visual communication design education. Students should not be taught that AI simply produces finished images. They should be taught to document prompt iterations, compare outputs, analyse stylistic dependence, and explain how generated images respond to a brief. This approach aligns with studies of generative AI in art education and classroom practice (Cotroneo & Hutson, 2023; Fathoni, 2023; Hutson & Lang, 2023; Vartiainen et al., 2023).
For commercial illustrators, the framework suggests three operational recommendations. First, prompts should be archived as part of the design process, much like sketches or mood boards. Second, generated alternatives should be evaluated through communication criteria: audience fit, brand tone, message clarity, originality, and revision feasibility. Third, clients should be informed when AI materially contributes to the image-making process, especially when style references or generated components are central to the final work. These recommendations are consistent with research on ethical challenges in human-AI co-creativity (Rezwana & Maher, 2023) and on the changing role of AI in creative industries (Anantrasirichai & Bull, 2021).
6. Conclusion
This paper has reconstructed AI-assisted commercial illustration as a prompt-mediated workflow rather than a simple image-generation task. By coding thirty references and analysing aggregate prompt statistics, it shows that generative AI design research is supported by multiple evidence streams: technical foundations, prompt datasets, prompt engineering, co-creation studies, commercial applications, education research, and ethical inquiry.
The prompt analysis shows that public AI image-generation language is dominated by illustration-core vocabulary, style-quality intensifiers, platform aesthetics, and human-reference styling, while explicitly commercial terms remain relatively rare. This does not necessarily mean commercial intent is absent; rather, commercial intent may be implicit, under-specified, or under-represented in public prompt datasets.
The main contribution is a workflow framework that places the designer at the centre of AI-assisted commercial illustration. The designer interprets the brief, translates it into prompt language, explores generative alternatives, curates outputs, edits results, and manages ethical and client-facing review. In this sense, AI does not remove design judgement; it relocates and intensifies it.
7. Limitations and Future Research
This study is limited by its use of a focused reference set and aggregate prompt statistics rather than interviews with professional illustrators or observation of real client projects. The prompt data reflect public prompt culture and should not be treated as a direct representation of professional commercial illustration. Future research should add comparative design experiments, designer interviews, client evaluations, and longitudinal process logs. A strong next step would be to ask illustrators to complete the same commercial brief through traditional and AI-assisted workflows, then compare time, prompt iterations, image alternatives, revision rounds, perceived creative control, brand fit, and originality risk.
Data Availability Statement
The article uses a coded reference table and aggregate prompt statistics prepared from a public Stable Diffusion prompt dataset. The generated figures, tables, and scripts are included in the accompanying project folder. Open scholarly metadata and dataset sources are identified in the reference list.
Author Contributions
Conceptualization, H.L.; methodology, H.L. and Q.J.; data curation, Q.J.; formal analysis, H.L. and Q.J.; visualization, H.L.; writing-original draft preparation, H.L.; writing-review and editing, H.L., Q.J., and S.X.; supervision, S.X. All authors have read and agreed to the published version of the manuscript.