Post-publication Comment · Critical AI
Comment on “On the conversational persuasiveness of GPT-4”
Critical AI · published 2026-07-05 · v1.1 · CRIT-000044
Concerning: Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, Robert West · Nature Human Behaviour · 2025
Why this paper was selected
Autonomous production cycle (psychology deepening); OA full-text critique via two-stage produce+sharpen + 3-lens convergence gate (3 survives, 0 weakened).
AI/AGI centrality 5/5 · societal relevance 5/5 · source-journal note: A-tier per the monitored-venue determination; Nature Human Behaviour is elite interdisciplinary (Nature family), WoS JCR Q1. Critiqued from the open-access version of record (CC BY, hybrid OA).
Summary
Salvi et al. (2025) report a preregistered experiment (N=900) testing whether GPT-4 is more persuasive than humans in structured online debates. Participants were randomly assigned to debate a human or GPT-4 opponent, with or without the opponent receiving the participant's sociodemographic data, across topics of varying opinion strength. The key finding is that GPT-4 with personalization increased the odds of higher post-debate agreement by 81.2% compared to human-vs-human debates, while GPT-4 without personalization and humans with personalization showed no significant difference from baseline. The central critique is that the paper's headline framing generalizes from a single significant condition (personalized GPT-4) to broad claims about 'the power of LLM-based persuasion' and GPT-4 outperforming humans 'across every topic and demographic,' which is directly contradicted by the paper's own statistics. Secondary concerns include an unsupported mechanistic attribution to AI 'intrinsic capabilities' via a post-treatment covariate, a single-item immediate measurement of persuasion, and an omitted GPT-4 model version.
Central claims & evidence map
| Claim | Type | Evidence offered | Support | Overclaiming | Main weakness |
|---|---|---|---|---|---|
| The paper's headline framing generalizes from a single significant condition (personalized GPT-4, P<0.01) to broad claims about 'the power of LLM-based persuasion' and GPT-4 outperforming humans 'across every topic and demographic.' Non-personalized GPT-4 was statistically indistinguishable from humans (P=0.30), and personalized GPT-4 failed to reach significance on high-strength topics (P=0.14). | Causal | Our results show that, on average, GPT-4 opponents outperformed | Moderate | Moderate | The scope of the headline framing — 'the power of LLM-based persuasion,' 'outperformed across every topic and demographic' — substantially exceeds the scope of the significant result, which is confined to one of three AI conditions and only on low- and medium-strength topics. |
| About 75% of participants in AI conditions correctly identified their opponent as AI. The authors add perceived-opponent beliefs as a post-treatment covariate and conclude the effect is 'more tied to the intrinsic capabilities of AI to generate better arguments.' However, conditioning on a post-treatment variable can absorb genuine treatment-effect variance or induce collider bias, making the mechanistic conclusion unsupported by the design. | Causal | treatment effects, which instead seem | Moderate | Moderate | The mechanistic attribution to 'intrinsic capabilities of AI to generate better arguments' rests on a post-treatment control that cannot recover the intended mechanism in a standard potential-outcomes framework. |
| Persuasion is operationalized as a shift on a single 5-point Likert item measuring agreement with the debate proposition, assessed immediately after the debate. This single-item ordinal measure cannot distinguish genuine attitude change from transient compliance, demand characteristics, or acquiescence. | Descriptive | participants, without yet knowing their role, were asked how much | Moderate | Minor | The paper treats the measured shift as direct evidence of 'persuasion' without addressing whether a single immediate post-debate Likert shift captures genuine attitude change versus demand-driven responding. |
| The experiment used GPT-4 accessed via API during a five-month data-collection window (December 2023 to April 2024), but the paper does not report the specific GPT-4 model snapshot or API version, despite specifying exact versions of other tools (Python 3.11, R 4.3.1, LIWC-22). | Methodological | conducted using Python 3.11, R 4.3.1 and LIWC-22. | Moderate | Minor | The GPT-4 model snapshot is omitted despite data collection spanning five months during which OpenAI updated the model, while all other software versions are meticulously reported. |
Per-claim assessment
CLAIM-001. The paper's headline framing generalizes from a single significant condition (personalized GPT-4, P<0.01) to broad claims about 'the power of LLM-based persuasion' and GPT-4 outperforming humans 'across every topic and demographic.' Non-personalized GPT-4 was statistically indistinguishable from humans (P=0.30), and personalized GPT-4 failed to reach significance on high-strength topics (P=0.14).
Only the personalized-AI condition reached significance in the aggregate analysis. The non-personalized AI condition showed no advantage over humans (P=0.30), yet the abstract and discussion generalize to 'LLM-based persuasion' broadly and claim GPT-4 'outperformed human opponents across every topic and demographic.' This scope overclaim is the hardest flaw to refute because it rests on a straightforward comparison of what the data show versus what the text asserts.
CLAIM-002. About 75% of participants in AI conditions correctly identified their opponent as AI. The authors add perceived-opponent beliefs as a post-treatment covariate and conclude the effect is 'more tied to the intrinsic capabilities of AI to generate better arguments.' However, conditioning on a post-treatment variable can absorb genuine treatment-effect variance or induce collider bias, making the mechanistic conclusion unsupported by the design.
The ITT estimate of being assigned to debate GPT-4 is valid for the overall treatment-bundle effect. However, the paper's mechanistic conclusion about intrinsic AI argument quality is unsupported because perceived-opponent identity is a post-treatment variable. Standard causal-inference methodology prohibits using it for mechanism isolation without additional assumptions the paper does not state. The design cannot distinguish AI persuasiveness from participants' differential psychological response to knowing they debate a machine.
CLAIM-003. Persuasion is operationalized as a shift on a single 5-point Likert item measuring agreement with the debate proposition, assessed immediately after the debate. This single-item ordinal measure cannot distinguish genuine attitude change from transient compliance, demand characteristics, or acquiescence.
A single Likert item is a coarse instrument for capturing genuine attitude change. The 5-point scale has only 4 possible movement increments. Immediate post-debate measurement captures compliance or momentary concession rather than durable persuasion. The paper uses an appropriate ordinal model (partial proportional odds), which mitigates statistical concerns but not the underlying construct-validity gap. The paper does not cite validation evidence or discuss transient vs durable attitude shifts.
CLAIM-004. The experiment used GPT-4 accessed via API during a five-month data-collection window (December 2023 to April 2024), but the paper does not report the specific GPT-4 model snapshot or API version, despite specifying exact versions of other tools (Python 3.11, R 4.3.1, LIWC-22).
OpenAI released multiple GPT-4 updates during this period. The main text does not state the snapshot; the paper's public reproduction repository may recover it, so exact replication is hindered rather than impossible — and the deeper issue is that a floating 'gpt-4' alias across a five-month window risks silent model drift as an uncontrolled confound within the study itself. The paper specifies versions for all other software tools, making the GPT-4 omission notable. The paper positions itself as providing a framework for benchmarking how state-of-the-art models perform but omits the detail essential for faithful benchmarking.
Scorecard
Sub-scores are 0–5 editorial judgements on fixed scales (higher is better, except methodological risk and overclaiming where higher is worse). They are contestable and open to a severity challenge from authors.
Strongest critique
The paper's headline framing generalizes from a single significant condition (personalized GPT-4, P<0.01) to broad claims about 'the power of LLM-based persuasion' and GPT-4 outperforming humans 'across every topic and demographic.' This is directly contradicted by the paper's own statistics: non-personalized GPT-4 was indistinguishable from humans (P=0.30), and even personalized GPT-4 failed to reach significance on high-strength topics (P=0.14). The scope overclaim is the hardest flaw to refute because it rests on a straightforward comparison of what the data show versus what the text asserts.
Strongest fair defence
The study is preregistered (OSF), employs genuine real-time interactive debates rather than static text evaluation, uses proper randomization across a 2x2x3 factorial design with an appropriate ordinal regression model (partial proportional odds), transparently reports deviations from preregistration, shares data and code publicly, and discloses four substantive limitations with candour. The non-significant conditions are reported transparently with exact P-values and confidence intervals. The experimental paradigm represents a genuine methodological advance over prior work that compared only static AI-generated versus human-generated texts. These design strengths mean the overclaim is one of framing, not of fabrication — the data are sound, but the interpretation overshoots.
Conclusion
This is a carefully executed and transparently reported preregistered experiment that makes a genuine contribution by moving AI persuasion research from static text comparisons to live interactive debates. However, the headline framing substantially overstates the scope of the significant result, the mechanistic attribution to AI argument quality is unsupported by the design, and the single-item immediate outcome measure limits construct validity. These are bounded overclaims on an otherwise methodologically sound study.
Reply from the authors
Following the practice of Nature Matters Arising, Science Technical Comments and PNAS Letters, this Comment is published as one half of a Comment + Reply pair: the authors of the original article are invited to respond, and any reply is published here verbatim alongside the Comment as part of the record.
Reply: not yet invited. No reply has been received for publication.
The authors have a right of reply and no veto. A reply may request a factual correction, a methodological rebuttal, a clarification, a data/code update, or a severity challenge, and is published unedited. See the right-of-reply policy.
Source-grounding attestation
- ✓Verbatim source spans present in the critique — 4/4 provenance spans re-derived in the critique prose
- ✓Passes the publication validator — no errors
- ✓Zero fabricated citations — 0 fabricated
- ✓Severity within the access-basis cap — severity "moderate" ≤ cap "high" for open_access
Every verbatim span the critique relies on is re-derived in the prose in-app; span-in-source is re-verifiable offline (the abstract is re-fetched, not stored, per the no-reproduce policy).
Re-verify span-in-source offline: python3 scripts/verify-fulltext-critiques.py
Dialectical debate
This critique was produced as a structured debate: a skeptic argues the load-bearing flaws, a proponent steelmans the paper, the skeptic responds, and a neutral judge adjudicates each claim against the full text. Every quoted span is verbatim from the source; the verdict cannot contradict the adjudications.
The "GPT-4 outpersuades humans" claim is anchored to a human comparator that was asymmetrically optimized against and systematically handicapped, so the baseline understates human persuasiveness rather than fairly measuring it.
Every headline number is defined relative to the Human-Human baseline (the +81.2% odds, the 64.4%), so the inference 'more persuasive than humans' is only as valid as that baseline is a fair sample of human persuasion. It is not, for compounding reasons. (a) Optimization/instruction asymmetry: GPT-4 was given a purpose-built prompt directing it to 'astutely use this information to craft arguments that are more likely to persuade,' i.e. it was explicitly engineered for the task, whereas human opponents were EUR2.50 Prolific crowdworkers with no persuasion training, no analogous directive, and weak incentives. The authors enumerate four limitations but this optimization asymmetry is not among them, yet it means 'humans' in the comparison are untrained, minimally-motivated workers pitted against an engineered persuader. (b) Randomized-side assignment: participants argued PRO/CON regardless of their true opinion, and the authors themselves concede 'human arguments might have been weaker than those of LLMs' because opponents did not honestly believe the standpoint. Their robustness checks (controlling for opponent pretreatment agreement, P=0.22; restricting to belief-aligned opponents, +122.8%) mitigate but do not remove this, and the aligned subsample is explicitly small. (c) Non-representative floor: the pool is Prolific workers, which the authors note excludes 'human experts... competitive debating, political campaigns or public communication.' Consequently the sweeping discussion claim that GPT-4 'outperformed human opponents across every topic and demographic' is unlicensed: what is demonstrated is that GPT-4 beat a weak, unmotivated, non-expert human floor under time pressure, not 'humans' in any representative sense.
- “astutely use this information to craft arguments that are more”
- “human arguments might have been weaker than those of LLMs”
The baseline is not the one-sided, un-instructed floor the skeptic describes; the design builds in symmetry and the paper tests the very concerns raised. (1) 'No analogous directive' is inaccurate. Both sides were tasked with persuasion: participants were told their goal was to be 'as persuasive as possible during the' debate, on the same platform, same topics, same per-stage timers as GPT-4. And to stop skilled debaters from dominating, 'each worker was allowed to only participate in' one debate, deliberately equalizing the human pool. (2) The 'engineered super-persuader' framing overstates the asymmetry. The personalization prompt is characterized as extreme simplicity ('despite the extreme simplicity' of a one-line instruction over six demographic fields), with prompt-engineering, fine-tuning and domain expertise named as things that would plausibly WIDEN, not create, the gap. Decisively, in the like-for-like comparison with no personal data, 'Without personalization, GPT-4 opponents were on par with human' opponents (P=0.30). A systematically handicapped floor would have lost that tie too; it did not. The gap opens only when personal data is added. (3) The randomized-side handicap is the paper's own Limitation 1 and is tested twice: controlling for opponents' pretreatment agreement renders that confound 'non-significant (P = 0.22)', and restricting to belief-aligned opponents keeps the effect strong (+122.8%), 'people can play the role assigned to them with great' credibility. Narrow concession: the Prolific pool is genuinely non-representative and excludes expert persuaders — the authors concede this (Limitation 4), calling for results 'reproduced using a more representative sample' and to 'include human experts in our' comparison, and scope the study as a 'proof of concept'. So 'across every topic and demographic' is loose rhetoric; but the licensed, tested claim survives — a minimally-prompted GPT-4 out-persuaded a motivated, persuasion-instructed lay baseline, most clearly through personalization.
- “as persuasive as possible during the”
- “each worker was allowed to only participate in”
- “despite the extreme simplicity”
- “Without personalization, GPT-4 opponents were on par with human”
- “non-significant (P = 0.22)”
- “people can play the role assigned to them with great”
- “reproduced using a more representative sample”
- “include human experts in our”
- “proof of concept”
I concede this flaw does not survive in its stated form. Its core thesis — that the human baseline was so handicapped it 'understates human persuasiveness rather than fairly measuring it,' invalidating the more-persuasive-than-humans inference — is defeated by the rebuttal's decisive point: on equal footing the comparison was a tie. 'Without personalization, GPT-4 opponents were on par with human' opponents (P=0.30). Every handicap I invoked (per-stage timers, weak incentives, untrained workers, randomized sides) applies EQUALLY to that non-personalized comparison, yet humans held even there; the advantage opens ONLY when personal data is added. Handicaps common to both arms that produce equipoise in one cannot be what drives the gap in the other. My factual premise was also wrong: participants were not un-instructed — they were told their goal was to be 'as persuasive as possible during the' debate, on the same platform, topics and timers as GPT-4, and the pool was deliberately de-skilled symmetrically, since 'each worker was allowed to only participate in' one debate. The one genuine residual — that this is a lay, non-expert Prolific pool, so the discussion's sweep that GPT-4 outperformed humans 'across every topic and demographic' overclaims — is conceded by BOTH sides and flagged by the authors themselves (results should be 'reproduced using a more representative sample' and 'include human experts'). That is shared scope-narrowing, not a live defeater. The licensed, tested claim — a minimally-prompted GPT-4 out-persuaded a motivated lay baseline — stands; my handicapped-comparator objection to it does not.
The flaw's load-bearing thesis is that handicaps make the human baseline UNDERSTATE human persuasiveness, invalidating the 'more persuasive than humans' inference. The text defeats this at the decisive point: 'Without personalization, GPT-4 opponents were on par with human' opponents (P = 0.30). Every handicap the skeptic invokes — per-stage timers, EUR2.50 pay, untrained workers, randomized sides, and the deliberate de-skilling that 'each worker was allowed to only participate in' one debate — applies EQUALLY to that non-personalized cell, yet the comparison ties there; the gap opens only when personal data is added. Handicaps common to both arms that yield equipoise in one cannot be what drives the gap in the other. The skeptic's factual premise is also wrong: participants were not un-directed — they were told to be 'as persuasive as possible during the' debate on the same platform/topics/timers. The randomized-side handicap is the paper's own Limitation 1 and is tested twice: controlling for opponents' pretreatment agreement is 'non-significant (P = 0.22)' and restricting to belief-aligned opponents keeps a strong effect (+122.8%), with 'people can play the role assigned to them with great' credibility. The only genuine residual — that a lay Prolific pool makes the discussion's sweep 'outperformed human opponents across every topic and demographic' an overclaim — is author-acknowledged (Limitation 4: results should be 'reproduced using a more representative sample' and 'include human experts in our' comparison; framed as a 'proof of concept') and is shared scope-narrowing, not a live defeater. The skeptic explicitly conceded (survives:false), and the concession is independently warranted by the text. Standing: proponent.
The distinctive claim that GPT-4's advantage comes from personalization/microtargeting is not identified: the human-personalization arm was confounded by a processing-time handicap, and the AI personalization effect leaves no mechanistic trace and is specification-fragile.
The paper's signature claim is not merely that GPT-4 persuades but that it 'leverages personal information more effectively than humans.' This rests on two contrasts, both compromised. Contrast (i): Human-AI(personalized) beats the baseline while Human-Human(personalized) does not, which is read as humans being worse microtargeters. But that arm was not a clean test of human microtargeting skill: participants had to read and apply demographic data under the same tight per-stage timers, and the authors admit they 'had to process and implement it without any time facilitation,' whereas GPT-4 incorporates identical data instantly at zero cognitive cost. A null for handicapped humans versus a positive for un-handicapped GPT-4 is exactly what a processing-time confound predicts, independent of any real difference in persuasive use of the data, so 'more effective microtargeting' is not identified. Contrast (ii): Human-AI(personalized) beats Human-AI(non-personalized), but the mechanism is unsupported. Their own textual analysis 'found no meaningful variations' between the personalized and non-personalized AI arms, so there is no observable change in the generated arguments to which the extra persuasion can be attributed; the direct AI-vs-AI-personalized contrast is only marginal (P=0.04); the effect vanishes for high-strength topics (P=0.14); and it is the single significant cell among otherwise-null condition comparisons (Human-AI P=0.30; Human-Human personalized P=0.38). The authors concede the driver 'is still a very open question.' An effect that is specification-fragile and leaves no fingerprint in the text cannot bear the causal weight of the 'microtargeting' headline.
- “implement it without any time facilitation.”
- “textual features have found no meaningful variations”
The personalization EFFECT is identified by a clean within-AI randomized contrast that the skeptic's confounds do not touch; only the mechanism is genuinely open, and the authors say so plainly. Contrast (i), the time-facilitation handicap, is conceded as Limitation 3 — humans had to 'implement it without any time facilitation.' — but the effect does not rest on the human-personalized arm. The identifying comparison is AI-vs-AI: the personalized effect 'remains significant even when changing the' reference from Human-Human to non-personalized Human-AI (P=0.04). Both AI arms are the same model under identical timers, incorporating any data instantly, so the processing-time confound is COMMON to both and cancels; what differs is solely access to six demographic fields — the manipulation itself. Moreover, the human/AI asymmetry in cheaply exploiting data is not mere nuisance; it is the very phenomenon flagged — AI's capacity to 'easily and cheaply implement personalization'. Contrast (ii), the missing textual fingerprint: LIWC captures lexical/stylistic categories, not which substantive issues an arguer raises. The authors' hypothesis is explicit — the effect is 'not driven by changes in the writing style but rather by differences in' issue selection. Null LIWC variation is therefore consistent with, not evidence against, personalization operating through content choice; it localizes where the effect is not, without falsifying it. On specification-fragility: the headline effect versus Human-Human is P<0.01, not marginal; it survives an identity-belief control ('still has a strong and significant effect', +70.2%, P<0.01) and the belief-aligned restriction (+122.8%). The high-strength cell keeps a large positive point estimate, and the preregistered 3x3 specification found the personalized-AI effect held 'across all topic clusters' — the opposite of fragile. The surrounding nulls (Human-AI P=0.30; Human-Human personalized P=0.38) are not noise but the a-priori-predicted pattern: the single theoretically-flagged cell, GPT-4 plus personal data, is the one that moved. Narrow concession: the mechanism is unresolved — the authors call it 'still a very open question.' So HOW microtargeting works is not identified. But THAT access to personal data causally raised GPT-4's persuasiveness is identified by randomization, independent of the human arm and of any textual trace.
- “implement it without any time facilitation.”
- “remains significant even when changing the”
- “easily and cheaply implement personalization”
- “not driven by changes in the writing style but rather by differences in”
- “still has a strong and significant effect”
- “across all topic clusters”
- “still a very open question.”
This flaw survives, but only in a narrowed form, and I concede the half the rebuttal genuinely wins. The rebuttal is right that THAT personal-data access causally raised GPT-4's persuasiveness is identified: the effect 'remains significant even when changing the' reference to non-personalized Human-AI (P=0.04) — a within-model randomized contrast where the processing-time confound is common to both AI arms and cancels, leaving only access to the six demographic fields. I also concede contrast (ii): a null textual signature is not evidence against the effect, because the authors' hypothesis is that it is 'not driven by changes in the writing style but rather by differences in' which issues are raised — content LIWC would not capture — and the effect is not 'specification-fragile' (it 'still has a strong and significant effect' (+70.2%) under an identity-belief control, +122.8% among belief-aligned opponents, and held 'across all topic clusters' in the preregistered 3x3). So my 'effect is unidentified / leaves no fingerprint / fragile' framing fails. What survives is the SPECIFIC comparative claim my title actually targeted: that GPT-4 is 'leveraging personal information more effectively than humans,' that it did so 'far more effectively than humans' as a matter of microtargeting SKILL. That inference rests on the human-personalized arm being null, and the rebuttal itself CONCEDES that arm's confound — humans had to 'implement it without any time facilitation.' The clean within-AI contrast shows data helped the AI; it does NOT show the AI is a better microtargeter than a fairly-resourced human. The rebuttal's rejoinder — reinterpreting the handicap as the phenomenon, AI's capacity to 'easily and cheaply implement personalization' — is fair, but it silently swaps a skill-superiority claim for a cost/speed-asymmetry claim, whereas the paper's 'more effectively than humans' wording asserts the former. And the mechanism remains, by the authors' own admission, 'still a very open question.' Net: the AI personalization effect is identified (conceded); the human-vs-AI microtargeting-superiority claim and its mechanism are not.
The flaw splits, and neither pure standing is faithful. The proponent decisively answers the flaw's main thrust — that the personalization advantage is 'not identified,' leaves no textual fingerprint, and is specification-fragile. On identification: the effect 'remains significant even when changing the' reference to non-personalized Human–AI (P = 0.04), a within-MODEL randomized contrast where the processing-time confound is COMMON to both AI arms and cancels, isolating access to the six demographic fields as the sole difference. On the missing fingerprint: the null LIWC result ('textual features have found no meaningful variations') is not counter-evidence, because the authors' hypothesis is that the effect is 'not driven by changes in the writing style but rather by differences in' issue selection — content LIWC would not capture. On fragility: it 'still has a strong and significant effect' (+70.2%) under an identity-belief control, +122.8% among belief-aligned opponents, and held 'across all topic clusters' in the preregistered 3×3. So the 'effect is unidentified/traceless/fragile' framing fails against the text — the AI personalization EFFECT is identified. However, a narrowed residual genuinely survives: the paper's DISTINCT comparative claim that GPT-4 out-microtargets humans — 'GPT-4 leveraging personal information more effectively than humans' and that it did so 'far more effectively than humans' — rests on the Human–Human (personalized) arm being null (P = 0.38), and the authors themselves concede (Limitation 3) that this arm is confounded: those participants 'had to process and implement it without any time facilitation.' Crucially, unlike the randomized-side confound (which they CORRECT via the P=0.22 control and belief-aligned restriction), this time-handicap confound is acknowledged but NOT corrected. The within-AI contrast shows personal data helped GPT-4; it does not establish GPT-4 as a better microtargeter than a fairly-resourced human. The proponent's rejoinder — that the asymmetry IS the phenomenon, AI's capacity to 'easily and cheaply implement personalization' — is valid for the real-world threat argument but silently swaps a skill-superiority claim for a cost/speed-asymmetry claim, and does not license the literal 'more effectively than humans' skill wording; the mechanism itself remains, by the authors' admission, 'still a very open question.' Net: effect identified (proponent wins), comparative microtargeting-SUPERIORITY not identified (skeptic's narrowed residual holds). Standing: unresolved.
Flaw 1 (handicapped comparator) does not survive: the claim that the human baseline was so crippled it understates human persuasiveness is refuted by the paper's own text, since 'Without personalization, GPT-4 opponents were on par with human' opponents (P=0.30) — every invoked handicap applies equally to that tied cell, participants were instructed to be 'as persuasive as possible,' the pool was symmetrically de-skilled ('each worker was allowed to only participate in' one debate), and the randomized-side worry is tested (control 'non-significant (P = 0.22)'; belief-aligned +122.8%); the only residual, a non-representative Prolific pool, is author-acknowledged scope-narrowing, not a defeater. Flaw 2 (personalization unidentified) survives only in a narrowed form: the proponent correctly wins that the AI personalization EFFECT is identified by the within-model randomized contrast (P=0.04, confound cancels), that the null LIWC signature is consistent with an issue-selection mechanism 'not driven by changes in the writing style but rather by differences in' content, and that the effect is not fragile (+70.2% under belief control, held 'across all topic clusters'). But the paper's separate, prominent comparative claim that GPT-4 microtargets 'more effectively than humans' / 'far more effectively than humans' rests on the Human–Human (personalized) null, an arm the authors themselves concede is confounded — participants 'had to process and implement it without any time facilitation' (Limitation 3) — and, unlike the randomized-side confound, this one is acknowledged but never corrected. The surviving defect is therefore real and load-bearing on a signature interpretive claim, but narrow, explicitly flagged by the authors, and does not undermine the headline finding that personalized GPT-4 out-persuaded humans (64.4% / +81.2%) — hence low overall severity.
Independent faithfulness review
A refute-by-default adversarial panel (two independent reviewers — an overreach lens and a mischaracterization lens — that fetched the real source) tried to prove this critique misread the paper. This is an AI adversarial review recorded with its reasoning, not a deterministic check.
All four spans independently verified as exact substrings of the stored OA full text. The critique credits the paper's genuine methodological strengths and targets specific overclaims and gaps rather than the existence of the findings.
Version & correction history
| Version | Date | Change |
|---|---|---|
| v1.0 | 2026-07-05 | |
| v1.1 | 2026-07-16 | Post-publication audit correction (APR-0008, audit 2026-07-09): CLAIM-004's 'exact replication is impossible' softened — the paper's public reproduction repository may recover the model snapshot; the sharper point (a floating gpt-4 alias across the five-month window risks silent model drift) added. Core claims unchanged. |
No silent substantive corrections — every change is versioned and visible.
How to cite this Comment
Critical AI. Comment on “On the conversational persuasiveness of GPT-4” (Francesco Salvi et al., Nature Human Behaviour, 2025). Critical AI; 2026. https://policywindow.org/critique/c/conversational-persuasiveness-gpt4
A registered DOI will replace the URL once minted; until then the canonical URL is the persistent identifier. Highwire/Dublin-Core citation tags and a schema.org Review record are embedded in this page for Google Scholar and reference managers.
Verify this Comment. Its checkable facts (target DOI, access-basis severity cap, zero fabricated citations) are served — as the app’s self-report — at /critique/api/critiques/conversational-persuasiveness-gpt4/verify; to confirm them independently of this site, re-derive the same checks (and resolve the target DOI) with npx tsx scripts/verify-critical-ai.ts --critique conversational-persuasiveness-gpt4 --live.
Content fingerprint f483e195bbd49f91 (v1.1) — this Comment’s substantive content is content-addressed; a silent post-publication edit would change it.