Published on: 25 September, 2026
After the Journal of Business Ethics has been delisted from FT50 list and our first article on this journal’s quality lapse, today we are sad again to report another article from the same journal. In addition to former FT 50 list, JBE is indexed/abstracted in Clarivate's SSCI (IF 2025: 6.3), Scopus (Q1), ABDC, ABS, PsycINFO, and many other so-called reputed indexing/abstracting outlets. For open access articles, JBE charges a hefty fee of USD 3,990.
The title of the article is “Generative Artificial Intelligence and Ethicality in Entrepreneurs' Creativity” by Mostafiz, Gali, Ahmed, Hughes, and Simeonova, published in JBE in March 2026 – a paper that promises to reveal the "dark side" of GenAI, but instead reveals the dark side of peer review and editorial oversight in JBE.
Our comprehensive forensic audit identified 14 classified contradictions, internal tensions, theory-measurement mismatches, post-hoc explanations, and reporting concerns. Let's dissect the most egregious flaws—starting with the foundational category error that invalidates the entire empirical exercise.
The paper's entire theoretical architecture rests on a single, unproven assumption: that using a tool is equivalent to engaging with it in the theoretically specified manner. This is the fatal logical fallacy at the heart of the study.
In Measurement subsection under Research Methods, the paper states: "We measured G-AI usage with nine items adopted from Sundaram et al. (2007)."
The Sundaram et al. (2007) scale was developed to measure salesforce automation (SFA) and customer relationship management (CRM) tools—specifically, technologies like Salesforce.com and Siebel SFA designed to automate customer data collection, manage contacts, and submit call reports. The outcomes it predicted were administrative performance and salesperson performance—routine, task-completion metrics in a structured sales environment.
The scale's three dimensions were (sample items reported):
· Frequency of use: "On average, how frequently have you been using [the] technology for your work?"
· Routinization: "My use of [the] technology has been incorporated into my regular work schedule."
· Infusion: "I am using [the] technology to its fullest potential for supporting my work."
The authors took a scale designed for structured, routine, task-completion tools (SFA/CRM) and imposed it on generative, creative, open-ended systems (ChatGPT, Gemini, Claude, DALL-E, Midjourney).
| Dimension | Sundaram et al. (2007) Context | Mostafiz et al. (2026) Context |
|---|---|---|
| Technology | SFA/CRM: structured data entry, contact management, report generation | G-AI: open-ended dialogue, creative generation, strategic ideation |
| User relationship | Passive tool adoption | Active co-creation partner |
| Outcome | Administrative/sales performance (efficiency) | Creativity and unethical creativity (cognitive/ethical) |
| Task nature | Routine, repetitive, well-defined | Novel, ambiguous, generative |
A scale that measures how often a salesperson logs a customer call cannot measure how creatively an entrepreneur engages with a generative AI system.
The paper states the scale was "adopted" from Sundaram et al. (2007). This is scientifically inaccurate. Adoption means using an existing scale without modification for the same construct in a similar context. Adaptation means modifying an existing scale for a new construct, new context, or new population—which requires re-validation of content validity, construct validity, and measurement equivalence.
The authors:
· Changed the referent from "technology" to "G-AI technology" with examples (ChatGPT, Gemini, Claude, DALL-E, Midjourney) that did not exist in 2007.
· Changed the outcome from administrative/sales performance to creativity and unethical creativity.
· Changed the population from sales personnel to entrepreneurs.
· Changed the context from structured sales environments to open-ended entrepreneurial ideation.
This is adaptation, not adoption. Yet the authors report none of the required psychometric re-validation: no expert panel for content validity, no cognitive interviews, no pilot testing, no exploratory factor analysis, no confirmatory factor analysis for the new construct structure. They simply assumed that a scale designed for 2007-era CRM software would seamlessly measure 2026-era generative AI.
The paper theorizes G-AI as: "an active participant in the creative process", a "creativity partner", and a mechanism for "attributional ambiguity". Yet the empirical measure asks only:
· "On average, how frequently have you been using G-AI technology for your work?"
· "My use of G-AI technology is a normal part of my work."
· "I am using G-AI technology to its fullest potential for supporting my work."
· Entrepreneur B uses G-AI weekly for strategic brainstorming, generating deceptive marketing copy, and finding regulatory loopholes. Score: Moderate frequency. Moderate routinisation.
Under the authors' scale, Entrepreneur A scores higher on "G-AI usage" than Entrepreneur B. Yet the paper predicts that higher usage leads to unethical creativity. Entrepreneur A—the spell-checker—is statistically predicted to be more unethically creative than Entrepreneur B, the strategic manipulator. This is nonsensical. The scale cannot distinguish a responsible administrative user from a malicious strategic exploiter. Frequency of use is not a proxy for the nature of use.
Three items measure "infusion," asking respondents to self-assess whether they are using G-AI "to its fullest potential" and "in the best fashion". These items do not measure use—they measure self-efficacy, overconfidence, or perceived mastery.
When narcissism is a moderator—a trait characterized by grandiose self-perception and overestimation of ability—the usage score becomes confounded with the moderator itself. A narcissistic entrepreneur is predisposed to rate their own G-AI mastery more highly, inflating their "usage" score. The predictor and the moderator are not independent. The model is structurally biased.
The scale includes zero items measuring:
· Prompt design or refinement (the actual act of directing G-AI).
· Iterative dialogue or revision (the co-creation process).
· Verification or fact-checking (the responsible use they mention).
· Task purpose (administrative vs. analytical vs. creative vs. deceptive).
· Cognitive delegation or reliance (the attributional blurring they theorize).
A scale that does not measure the theoretical mechanism cannot test the theory. The authors are inferring a psychological process—attributional reorientation, moral distancing, self-regulation depletion—from a behavioral frequency count. This is not just a measurement weakness; it is a logical invalidation of the entire empirical exercise.
The paper reports in the “Assessment of potential biases” section: "The results (χ² = 1239.8472, df = 148, CMIN/df = 8.377, RMSEA = 0.478, CFI = 0.184) differed markedly from the six-factor confirmatory factor model (χ² = 421.397, df = 238, CMIN/df = 1.77, RMSEA = 0.049, CFI = 0.901)."
However, the Online Supplementary Table 2 reports a different measurement model for the same pooled data: "χ² = 154.834, df = 112, CMIN = 1.382, CFI = 0.918, RMSEA = 0.048."
The supplementary table lists "CMIN = 1.382." However, dividing the reported chi-square χ² = 154.834 by the degrees of freedom df = 112 yields1.3824. This indicates that 1.382 is actually the CMIN/df ratio, not the CMIN (χ²) value itself.
Which model is correct? The paper does not reconcile this discrepancy. The reported χ² differs by 266 points, with different degrees of freedom (238 vs. 112). This is not a minor rounding error—it is a fundamental reporting inconsistency. Either the main text or the supplement contains an incorrect measurement model. Editors and reviewers apparently failed to notice this basic discrepancy.
The authors claim (see online supplementary file): "The findings suggest that the items and the constructs maintain consistent reliability across the samples (also enabling the study to exclude country as a dummy variable), validating the measures for hypothesis testing."
Yet their own data show:
| Metric | Recommended Threshold | Observed Range | Threshold Status |
|---|---|---|---|
| ΔCFI | ≤ 0.0100 | 0.0260 – 0.0380 | Exceeded across all models (Metric, Scalar, Partial) |
| ΔMcNCI | ≤ 0.0200 | 0.0038 – 0.0248 | Exceeded at the Scalar level (0.0248) |
| Δγˆ | ≤ 0.0010 | 0.0121 – 0.0289 | Exceeded across all models (Metric, Scalar, Partial) |
The paper's ΔCFI values (0.026–0.038) exceed the 0.01 threshold. The ΔMcNCI (0.0248) exceeds the 0.02 threshold. The Δgamma hat (0.0289) exceeds the 0.001 threshold.
The authors claim invariance while their own data show non-invariance. This is a direct empirical contradiction. Pooling data from two culturally different countries (India and Bangladesh) when measurement invariance is not established is methodologically indefensible. Their own statistical tests disprove their central assumption—yet they proceed to pool the data anyway.
The paper measures "unethical creativity" with items such as:
"I develop innovative ways to skirt ethical rules."
"I engage in creative adjustment of data."
"I creatively employ workarounds to ethical policies."
"I creatively take advantage of loopholes."
Respondents are asked to self-report engaging in unethical acts. This is a severe social desirability problem. Even with supervisor ratings (collected at t2), the entrepreneurs introduced the managers who rated them—potentially selecting favorable raters. The authors do not report inter-rater agreement or test for social desirability.
Hypothesis 1: The paper states: "G-AI usage positively influences creativity."
Result: "G-AI usage has no significant effect on creativity (β = 0.029, p > 0.05, t = 0.72); therefore, H1 is not supported."
The authors then offer multiple post-hoc explanations in the Discussion:
· "Not all entrepreneurs consider G-AI as a creativity-enhancing tool."
· "Recombination does not automatically translate into higher-quality creativity."
· "A ceiling-effect explanation is plausible."
· "G-AI may primarily function as a productivity-enhancing tool rather than a driver of fundamentally new ideas."
These explanations are plausible—but they are post-hoc storytelling, not theory-driven predictions. The original hypothesis argued that attributional ambiguity would inflate internal attributions and boost creativity. When the hypothesis failed, the authors invented new mechanisms. This is not how theory-testing works.
The paper effectively shifts its theoretical position after seeing the data—a classic case of HARKing (Hypothesizing After Results are Known).
The paper claims to advance "attribution theory" and advance "self-regulation theory." The abstract states: "Our study advances attribution theory and self-regulation by revealing how entrepreneurs' internal conflicts, such as the tension between aspirational creativity and opportunistic self-interest, interact with their self-control to produce unethical creativity."
Yet the paper argues that G-AI usage leads to unethical creativity through:
· Attributional ambiguity (not measured).
· Moral distancing (not measured).
· Lower perceived accountability (not measured).
· Self-regulation failure (not measured).
No mediating variable is measured. The paper infers a psychological process from a correlation between G-AI usage and unethical creativity, moderated by personality traits. This is an untenable leap. Correlation is not causation, and association is not mechanism.
The authors claim to advance attribution theory and self-regulation theory. Yet they measure neither attributional processes nor self-regulation. They measure personality traits and behavioral outcomes—and infer the mechanisms.
A theoretical contribution requires testing the mechanism. This paper tests only the endpoints.
Under sub-section "Sample and data collection", the authors state: "To ensure adequate statistical power and external validity, we calculated the minimum sample size using Cochran's (1942) formula with a 95% confidence level, ±7% margin of error, and maximum variability (p = 0.50). For Bangladesh, the required minimum sample was 179; for India, the required minimum was 192."
While the calculation of these numbers is mathematically sound, describing them as ensuring "adequate statistical power" for complex regression or moderation models is technically a misapplication. Cochran's formula calculates precision for estimating descriptive proportions, not statistical power for testing multivariate hypotheses or interaction effects. Reviewers or methodologists who look closely at the terminology might flag the conflation of survey precision with statistical power, even though the math checks out. The study is likely underpowered to detect interaction effects, which may explain why H4 (psychopathy) failed. This is a basic methodological oversight that should have been caught during review.
Hypothesis 4: The paper states: "The relationship between G-AI usage and unethical creativity is moderated by psychopathy, such that the relationship...will be strongest for entrepreneurs who exhibit high levels of psychopathy."
Result: "the moderating effect of psychopathy is non-significant (β = 0.003, p > 0.05, t = 0.33)."
The authors offer two speculative explanations in the Discussion:
1. "Psychopathy is invariably maladaptive (Smith & Lilienfeld, 2013), while innovation-centered activities demand adaptability skills (Smith & Webster, 2018)."
2. "The need for immediate gratification and impulsivity for psychopaths is less likely to be associated with liking a technology that would bring power over others in the long run rather than immediately."
The second explanation introduces a new temporal mechanism (immediate vs. delayed gratification) that was never mentioned in the theory section. Neither gratification preference nor perceived timing of benefits was measured. The authors are inventing explanations for why their hypothesis failed—without evidence.
The abstract states: "Our study advances attribution theory and self-regulation by revealing how entrepreneurs' internal conflicts, such as the tension between aspirational creativity and opportunistic self-interest, interact with their self-control to produce unethical creativity."
Neither internal conflict, nor self-control, nor the interaction between them, was measured. The abstract presents inferred mechanisms as empirical findings. This is not just rhetorical flourish—it misrepresents what the study actually did. The strongest summary uses causal language beyond the observational design.
G-AI usage is self-reported by individual entrepreneurs. Yet the discussion frequently generalizes to firm development, management, and organizational governance. Individual use and organizational adoption are not equivalent. Conclusions cross levels without corresponding multilevel evidence.
The paper states: "We measured G-AI usage with nine items adopted from Sundaram et al. (2007)."
However, Appendix I lists:
· Two frequency items
· Three routinisation items
· Three infusion items (Sundaram et al. (2007) mentioned 4 items)
These are totals eight items, not nine. The paper reports nine items but the appendix contains eight. The authors missed one item of infusion scale. On page 105, Sundaram et al. (2007) mentioned “Jones et al. (2002) developed the scale for infusion, and we developed the measure of routinization for this study.” However, Mostafiz et al. (2026) mentioned that they adopted the scale from Sundaram et al. (2007). This is a clear dishonest reporting as the infusion scale was developed by Jones et al. (2002). This is an instance of dishonest reporting that editors and reviewers should have caught.
The paper uses Enron's accounting misconduct and Microsoft's Tay chatbot as examples of unethical creativity (see Introduction, second para). However:
· Enron predates G-AI and does not substantiate the human-AI mechanism.
· Tay was a chatbot—not comparable to contemporary generative AI.
The examples illustrate the outcome (unethical behavior) but not the causal process (attributional displacement via G-AI). The constructs do not fit.
To test for endogeneity, the authors included "entrepreneurial orientation" (EO) as a missing variable. However:
· They do not report how EO was measured.
· They do not state whether EO was collected at t1 or t2.
· They do not provide the EO scale or its psychometric properties.
Simply adding Entrepreneurial Orientation (EO) as a control variable and showing that the results don't change is a robustness check for omitted variable bias, not an econometric test for endogeneity. Endogeneity typically arises from simultaneous causality, measurement error, or unobserved omitted variables. Adding a single observed variable (EO) does not resolve endogeneity or prove it doesn't exist. Reviewers should have asked the authors to conduct formal econometric remedies for endogeneity, such as Instrumental Variables (IV/2SLS), a Lewbel’s heteroskedasticity-based instrument approach, or a Hausman specification test.
The authors report a Heckman two-stage test with no significant Inverse Mills Ratio. However:
· The selection equation is not clearly described.
· The variables used to predict selection into the sample are not specified.
· Without a credible selection model, the Heckman correction is meaningless.
The Heckman two-stage model requires at least one exclusion restriction—an exogenous variable that affects the selection/participation decision (first stage) but does not directly affect the main outcome variable (second stage).
If one uses the exact same predictors in both stages without a valid exclusion restriction, the Inverse Mills Ratio (IMR) is identified purely through non-linearity (distributional assumptions), which econometricians heavily criticize as weak and prone to multicollinearity. Authors should have explicitly stated what instrument(s) drove the first-stage selection equation.
Now, one cannot help but ask: where were the distinguished editors and reviewers of the Journal of Business Ethics? Did they fail to notice that the empirical measure captures frequency of usage while the theory requires quality of creative engagement? Did they not compare the main text and the supplement to spot the blatantly different measurement models (χ² = 421 vs. 154)? Did they not apply Cheung and Rensvold's (2002) thresholds to the measurement invariance results—and realize that the authors' own data disprove their invariance claim? Did they not question how a 2007 CRM adoption scale could possibly measure 2026 generative AI co-creation? Did they not count the items in Appendix-I and notice the discrepancy between "nine items" and the eight actually listed? Did they not spot that the paper measures frequency but theorizes quality?
The peer review process at former FT 50 journals should be designed to catch exactly these foundational category errors. To miss a construct validity failure of this magnitude—where the scale measures SFA/CRM software adoption in 2007 but is deployed to measure generative AI co-creation in 2026—signals either:
1. A complete lack of methodological expertise among the reviewers.
2. A disturbing relaxation of editorial standards in pursuit of "novel" AI-themed papers.
3. Nescient and/or relaxed peer review practices that allow seriously flawed articles to slip through.
The paper reports conventional psychometric indicators (Cronbach's alpha, AVE, CFI) and significant regressions—and the editorial team appears to have accepted statistical significance as sufficient evidence of scientific validity. This is not how rigorous peer review works. Content validity, construct validity, and the fundamental alignment between theory and measurement should be the minimum requirements for publication in a former FT 50 journal.
The paper's 14 high-severity contradictions—including the foundational G-AI construct failure, the measurement model inconsistency, the measurement invariance violation, the catastrophic CMV model fit, the unmeasured theoretical mechanisms, the item count discrepancy, and the "adopted" vs. "adapted" misrepresentation—represent a systemic failure of the peer review process.
Beyond the purely statistical and methodological blunders, the publication of this paper carries significant real-world and downstream academic costs that necessitate its urgent retraction.
This paper threatens to corrupt the scientific record. Scholars conducting systematic reviews, meta-analyses, or theory-extension studies may treat its findings as a credible empirical anchor—a false anchor. Consider a future doctoral student building a framework on "G-AI usage and creativity." They might cite this paper as evidence that G-AI does not enhance creativity (H1 failure) and does enhance unethical creativity (H2 support), shaping their own hypotheses and research design. However, because the G-AI usage construct never measured creative engagement—only frequency of administrative use—any theory built on these findings rests on a foundation of sand. Worse, the post-hoc theorizing on attribution theory, unmoored from actual measurement, risks sending theoretical development down fruitless paths. Future researchers might waste years and funding investigating attributional processes that the paper inferred but never tested, normalizing HARKing (Hypothesizing After Results are Known) as an acceptable practice in top-tier journals. This erodes the self-correcting nature of science and rewards methodological negligence over intellectual rigor.
The managerial implications are dangerously misleading. The authors recommend "close monitoring" of entrepreneurs with high Machiavellianism and Narcissism when using G-AI. However, because their scale fails to distinguish benign administrative use (grammar checking, summarization) from strategic deceptive use (regulatory arbitrage, manipulative marketing), an organization might unjustly penalize a highly frequent, honest user (the email summarizer) while completely missing a low-frequency, high-impact cheater (the loophole exploiter). This could create a toxic culture of personality-based surveillance, exposing firms to lawsuits for discriminatory monitoring and eroding employee trust—all based on a statistically spurious correlation. Moreover, entrepreneurs reading this paper might be falsely reassured that high-frequency G-AI use does not affect their creativity, leading them to underinvest in training for meaningful human-AI co-creation. Alternatively, they might be wrongly stigmatized for using G-AI frequently, fearing it signals unethical tendencies. Either way, the paper provides no actionable, valid guidance.
The paper's call for "robust regulations on G-AI adaptation" is built on sand. If policymakers cite this study to justify strict curbs on general AI access—arguing that G-AI causes unethical creativity—they risk stifling the very innovation G-AI could foster in developing economies like Bangladesh and India. For instance, a government agency drafting AI ethics guidelines might use this paper as evidence that G-AI inherently facilitates cheating, leading to overly restrictive licensing requirements or outright bans on certain G-AI tools in entrepreneurial sectors. This would disproportionately harm small and medium enterprises that rely on G-AI for legitimate productivity gains, while completely ignoring the actual nuanced risks (e.g., data privacy, algorithmic bias, intellectual property infringement). Basing heavy compliance burdens on a construct that measures the quantity rather than the nature of G-AI use is policy malpractice. It diverts regulatory resources from real threats to phantom ones.
Consequently, we unequivocally demand a retraction. The foundational error—imposing a pre-GAI CRM scale onto a post-GAI co-creation context—deprives the paper of scientific validity. The measurement model inconsistency, invariance failures, catastrophic CMV results, and unmeasured mechanisms compound this invalidity. Retracting this paper is not merely correcting the record; it is a necessary act to:
1. Preserve the integrity of the Journal of Business Ethics and the FT 50 brand.
2. Protect practitioners from misguided, potentially harmful recommendations.
3. Ensure policymakers are not misled by pseudoscience masquerading as empirical rigor.
4. Safeguard future researchers from wasting resources on a faulty empirical foundation.
5. Reaffirm that construct validity—the alignment between theory and measurement—is non-negotiable, regardless of statistical significance.
Leaving this paper published rewards methodological negligence, penalizes honest scholars who adhere to rigorous standards, and erodes public trust in academic research. Springer Nature and the editors of JBE must act decisively to retract this article and issue a formal expression of concern.
We call upon the editors of the Journal of Business Ethics, Springer Nature, and the academic community to retract this paper immediately.
Don’t forget to leave a comment if you find this article interesting.
About the Author: This investigation was conducted by an anonymous professor at a renowned European university.
"Scholarly Criticism" is launched to serve as a watchdog on Business Research published in so-called Clarivate/Scopus indexed high quality Business Journals. It has been observed that, currently, this domain is empty and no one is serving to keep authors and publishers of journals on the right track who are conducting and publishing erroneous Business Research. To fill this gap, our organization serves as a key stakeholder of Business Research Publishing activities.
For invited lectures, trainings, interviews, and seminars, "Scholarly Criticism" can be contacted at Attention-Required@proton.me
Disclaimer: The content published on this website is for educational and informational purposes only. We are not against authors or journals but we only strive to highlight unethical and unscientific research reporting and publishing practices. We hope our efforts will significantly contribute to improving the quality control applied by Business Journals.