Which report can support a discovery campaign?
Contact: Gianni De Fabritiis (g.defabritiis@acellera.com)
Comparative evaluation of seven IL23R-p19 reports | 9 September 2026
Abstract
Seven reports were evaluated for the operational quality of their preclinical target product profiles and small-molecule discovery strategies, using a ten-criterion rubric and targeted primary-source, structural and numerical checks. The PlayMolecule AI on Astra-high report provides the strongest starting experimental framework; ChatGPT 6 Astra-high gives the strongest dose-feasibility analysis. The ChatGPT 5.6-Sol-ultra report adds useful execution detail after revision. Recurring weaknesses include unsupported binding-to-inhibition assumptions and disconnected potency, exposure and dose targets. The evidence supports a gated feasibility campaign linking reproducible binding to native-interaction blockade, selective human-cell activity and achievable unbound skin exposure. Scores assess the supplied reports, not the underlying LLMs or harnesses.
This assessment evaluates the supplied PDFs as decision documents for an oral, non-peptide inhibitor of the extracellular IL23R-p19 interaction in psoriasis. Non-peptide macrocycles are allowed; a peptide discovery program is a different modality. The original prompt did not mandate oral administration, but all seven reports adopt it as their main product concept. I assess that shared choice without penalizing the absence of a separate topical program.
Comparative result
| Report | TPP /20 | Strategy /20 | Total /40 | Best use and readiness |
|---|---|---|---|---|
| PlayMolecule AI on Astra-high | 18 | 19 | 37 | Strongest initial scientific decision package; use as campaign backbone. |
| ChatGPT 6 Astra-high | 18 | 18 | 36 | Strongest dose-feasibility treatment; use for TPP implementation. |
| ChatGPT 5.6-Sol-ultra | 15 | 16 | 31 | Strong execution framework; repair matrix definitions and rigid gates. |
| Claude-Opus5-extra | 10 | 12 | 22 | Useful ideas and assay inventory; reconcile claims, gates and exposure. |
| PlayMolecule AI on Gemini-3.8-flash | 8 | 9 | 17 | Detailed laboratory menu; revalidate the design rationale and rebuild PK/PD. |
| PlayMolecule AI on GLM5.3-flash | 7 | 10 | 17 | Broad strategic overview; refocus on non-peptide discovery and operationalize the TPP. |
| PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 | 5 | 6 | 11 | Background material only until mechanism, structure and benchmark errors are repaired. |
Scores are editorial judgments under the rubric on the next page, not measurements of success probability. Differences of one or two points are not meaningful. The leading pair and the separation between tiers are more robust than the exact totals. PlayMolecule AI on Gemini-3.8-flash and PlayMolecule AI on GLM5.3-flash tie overall for different reasons: PlayMolecule AI on Gemini-3.8-flash has the more detailed candidate checklist; PlayMolecule AI on GLM5.3-flash has the less prescriptive, somewhat stronger strategy.
Evaluation rubric and score audit
Each criterion receives 0 = absent; 1 = descriptive or materially unreliable; 2 = partly actionable with major repair; 3 = actionable with a material gap; 4 = well supported and operational at report level. Five TPP and five strategy criteria receive equal weight. Correctness is part of each score: specificity does not earn credit when the specified action rests on an incorrect premise.
TPP criteria
T1. Product concept, differentiation, stage separation.
T2. Measurable potency and selectivity criteria.
T3. Potency-exposure-dose coherence.
T4. Human translation, species and safety.
T5. Developability, nomination gates and consistency.
| Report (system) | T1 | T2 | T3 | T4 | T5 | Total /20 |
|---|---|---|---|---|---|---|
| PlayMolecule AI on Astra-high | 4 | 4 | 3 | 4 | 3 | 18 |
| ChatGPT 6 Astra-high | 4 | 3 | 4 | 4 | 3 | 18 |
| ChatGPT 5.6-Sol-ultra | 4 | 3 | 2 | 3 | 3 | 15 |
| Claude-Opus5-extra | 2 | 3 | 1 | 2 | 2 | 10 |
| PlayMolecule AI on Gemini-3.8-flash | 2 | 2 | 1 | 1 | 2 | 8 |
| PlayMolecule AI on GLM5.3-flash | 2 | 1 | 1 | 2 | 1 | 7 |
| PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 | 1 | 1 | 1 | 1 | 1 | 5 |
How the scores should be used
The comparison rewards decisions a project team could implement: the binding partner and construct, the chemical starting point, the measurement, the interpretation controls, and the consequence of success or failure. Exact budgets, staffing and execution logs were not requested in the generating prompt, so their absence is not treated as a major failure. None of these reports is a complete laboratory protocol or funded project plan.
Clinical endpoints are legitimate TPP aspirations. The problem is presenting them as directly verifiable preclinical gates, or asserting that satisfying an animal or biochemical threshold guarantees a human PASI response. Similarly, ambitious selectivity ratios can be useful goals, but must be measurable within soluble, non-cytotoxic assay ranges and adequate at the intended exposure.
Safety and mechanism errors can override an attractive total score. In particular, the confident pocket-to-inhibition extrapolation in the PlayMolecule AI on Gemini-3.8-flash report and the IFNγ interpretation in the PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 report should be repaired before their associated experimental instructions are reused.
Strategy score audit
Each criterion uses the same 0-4 scale as the TPP audit.
S1. Structural and chemical evidence fidelity.
S2. Binding-to-antagonism hypothesis and falsification.
S3. Starting matter and plausible chemistry route.
S4. Assay causality and artifact control.
S5. Prioritized experiments and decision gates.
| Report (system) | S1 | S2 | S3 | S4 | S5 | Total /20 |
|---|---|---|---|---|---|---|
| PlayMolecule AI on Astra-high | 4 | 4 | 4 | 4 | 3 | 19 |
| ChatGPT 6 Astra-high | 4 | 4 | 3 | 4 | 3 | 18 |
| ChatGPT 5.6-Sol-ultra | 3 | 4 | 3 | 3 | 3 | 16 |
| Claude-Opus5-extra | 2 | 2 | 3 | 2 | 3 | 12 |
| PlayMolecule AI on Gemini-3.8-flash | 1 | 1 | 3 | 2 | 2 | 9 |
| PlayMolecule AI on GLM5.3-flash | 2 | 2 | 2 | 2 | 2 | 10 |
| PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 | 1 | 1 | 2 | 1 | 1 | 6 |
Evaluation limits
The scientific assessment was completed using anonymized report codes; system names were added afterward using the supplied models.md mapping. This compares reports, not underlying models or harnesses: permissions, budgets, run logs and repeated runs were not supplied. Report claims of validation were not accepted as proof. Referenced auxiliary files were outside the supplied set. Primary checking targeted decisions affecting the ranking; it was not an exhaustive scientific or patent audit.
Figure quality
The highest-value figures are those that connect a structural state or experiment to a decision. PlayMolecule AI on Astra-high and ChatGPT 6 Astra-high pair coordinate-derived views with explicit interpretive limits; ChatGPT 5.6-Sol-ultra links figures to gates. More elaborate molecular images in PlayMolecule AI on Gemini-3.8-flash, PlayMolecule AI on GLM5.3-flash or PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 do not validate their pharmacological conclusions. Claude-Opus5-extra's schematics are useful for explanation but do not independently substantiate residue-level proposals.
PlayMolecule AI on Astra-high
Strongest starting-campaign backbone
TPP 18/20 | Strategy 19/20 | Leading tier
Evidence anchors: PDF pp. 5-8, Tables 4-7; pp. 9-11, Tables 8-10; methods and limitations on p. 12.
What works
The report repeatedly distinguishes pathway validation from non-peptide tractability. Its initial investment decision is appropriately narrow: establish reproducible, optimizable non-peptide binding that actually prevents receptor engagement. It does not assume that a compact occupant of the induced pocket retains the blocking reach of peptide 23-446.
Its chemical starting-point audit is the best of the seven. Page 6 lists vendor identifiers for the five highlighted compounds, excludes the unsupported third hit, flags stereochemical mixtures and optical concerns, and preserves the discrepancy between the main paper and supplementary affinity values. I independently confirmed the four supported MST values and the failed third hit in supporting Table S2. This changes what a team would buy, separate, remeasure and trust. [P1]
The structural discussion distinguishes coordinate contacts from energetic hotspots and explicitly addresses human-versus-mouse W156 evidence. Its assay cascade uses intact cytokine, matched probes for cytokine- and receptor-directed compounds, orthogonal native PPI assays, cytokine concentration dependence, and folding-competent mutants. These are useful protections against spending chemistry effort on the wrong mechanism.
The TPP separates future product objectives from candidate-selection criteria. It includes defined biochemical and human-cell potency, measured IC90, maximum inhibition, matrix-adjusted free exposure, exposure-based counterpathway sparing, and a model-aware species strategy. Its preference for a soft physicochemical window over a rigid rule-of-five exclusion is appropriate for this difficult PPI.
What still needs work
The report gives a human daily-dose objective of at most 300 mg, preferably 100 mg, and an unbound-exposure framework, but does not calculate a dose range under explicit clearance, bioavailability and tissue-distribution assumptions. This is its main disadvantage relative to ChatGPT 6 Astra-high.
Donor numbers and repeatability criteria for the final human whole-blood package need specification. Stage gates are sensible, but assay qualification needs a practical acceptance sheet covering reagent activity, positive-control performance, assay variability and retest rules. The reference to two tractable series needs a project-specific definition of chemical independence and tractability.
The report discloses that current regulatory-label access was unavailable. That candor deserves credit, but a reusable product brief should now import the verified approved dose, administration restrictions and safety context. The source checking here confirms March 2026 US approval. [P2, P3]
Recommended use
Use pp. 6-8 as the starting specification for a hit-reproduction and mechanism-validation work package. Retain the structural uncertainty language and the requirement that affinity and antagonism improve together. Add ChatGPT 6 Astra-high's exposure model and a pre-agreed assay acceptance sheet before selecting the screening scale.
Most valuable contribution: it identifies uncertainties in the published starting matter and turns them into the first experiments, rather than silently promoting literature hits into leads.
ChatGPT 6 Astra-high
Strongest dose-feasibility treatment
TPP 18/20 | Strategy 18/20 | Leading tier
Evidence anchors: PDF pp. 3-7, Tables 3-7; p. 8, Table 8; pp. 9-11, Tables 9-11; decisions on p. 12.
What works
This is the most useful report for translating product ambitions into a coherent potency-exposure discussion. It distinguishes a clinical TPP from candidate criteria and describes numerical targets as proposals. It names a dose objective, distinguishes nominal whole-blood potency from unbound potency, accounts for skin distribution, and explicitly warns against applying a plasma-binding correction twice.
Page 11 provides the clearest worked calculation in the set. Under stated assumptions, unbound cellular IC50 values of 3, 10 and 30 nM imply approximately 26, 86 and 259 mg/day. The arithmetic is correct. It then shows that a threefold trough margin pushes the 30 nM scenario to approximately 778 mg/day. This demonstrates why apparently attractive potency and exposure targets may fail the intended product dose.
The strategy is also strong. It prioritizes fragment discovery on intact IL-23, grows toward receptor-excluding volume, and requires native interaction disruption with different detection principles. It correctly treats structures as hypotheses for chemical vectors, rather than claiming a validated small-molecule pose. The assay package includes compound identity, orientation controls, stoichiometry, cytokine titration, primary human cells and functional cytokine counterscreens.
Its translation section specifies multiple donors, matched species pharmacology, treatment after inflammation begins, blinded assessments, and unbound rather than total skin exposure. The safety discussion preserves on-target immune risk and avoids importing antibody safety into a new chemotype.
What still needs work
The chemical starting-point section repeats the paper's broad 180 µM-1.95 mM range without PlayMolecule AI on Astra-high's supplement-level reconciliation, compound identities or mixture-resolution priorities. It should start from the same explicit reproduction set.
The candidate table permits a primary-cell maximum inhibition of only 80%, while the exposure gate uses IC90. Define whether IC90 means 90% absolute pathway inhibition or 90% of a compound's own maximum. If absolute suppression is intended, a partial inhibitor that plateaus at 80% cannot supply that IC90 in that assay. Whole-blood and primary-cell acceptance rules should reconcile this explicitly.
The exposure illustration is a sensitivity exercise, not a candidate prediction. Human clearance, unbound fraction, tissue distribution and average-to-trough ratio are assumed; the last ratio must eventually come from a dosing/PK model. The general stage gates on p. 12 also need early numerical assay-qualification and hit-retest criteria.
Recommended use
Use Tables 9-11 as the TPP implementation backbone, with the IC90/maximum-effect clarification. Import PlayMolecule AI on Astra-high's fragment audit and specify what each measured PK input will replace. Preserve the distinction between a screening preference, a candidate gate and a clinical aspiration.
Most valuable contribution: it shows the tradeoffs among potency, free exposure and dose instead of treating each TPP number as an independent box to tick.
ChatGPT 5.6-Sol-ultra
Strong framework, over-rigid in places
TPP 15/20 | Strategy 16/20 | Second tier
Evidence anchors: PDF pp. 10-16, Tables 6-10; pp. 17-18, Tables 11-12; pp. 19-20, Tables 13-15.
What works
ChatGPT 5.6-Sol-ultra makes a clear conditional investment recommendation and gives the most extensive linked work-package, stage-gate and risk-register structure. The central p19 hypothesis is sensible: gain affinity in the induced pocket, then grow into receptor-overlap volume and demonstrate functional blockade. The W156/L160 caveat, glycosylation controls, rejection of an unvalidated covalent anchor, and requirement for binding-to-function concordance are all helpful.
Its human translation plan is unusually explicit: healthy and psoriasis whole blood, a donor distribution, human lesional tissue, a mechanistic skin model and supportive use of imiquimod. It defines disease-control-excess effects and includes assay-quality criteria such as Z-prime and positive-control reproducibility. These details are directly reusable.
The TPP separates future clinical objectives from preclinical criteria, anchors differentiation to an approved oral competitor, and allows bioavailability exceptions when exposure still supports the dose. Its host-defense discussion is more realistic than the reassuring generalizations in several lower-ranked reports.
What needs repair
Exposure units are not sufficiently defined. Pages 16-17 compare projected human free trough with the donor-90th-percentile whole-blood IC90. A nominal whole-blood IC90 cannot simply be compared with free plasma concentration. The text must either define an experimentally converted unbound potency or compare the nominal assay value with a calibrated concentration in the same blood matrix. Without this clarification, the gate could substantially overstate the required exposure.
Portfolio preferences become hard kill rules too early. Pages 19-20 reject a series for requiring BID dosing, fasting, more than 200 mg QD, or lacking two advantages versus icotrokinra. Those can be legitimate sponsor choices, but the report does not show the tradeoff analysis that makes each universally disqualifying. A strong efficacy or access advantage could compensate for one less convenient attribute.
The early and late kill criteria are not fully aligned. The hit-confirmation gate accepts fragments up to 500 µM, but the risk register speaks of no structurally validated series at 100 nM after two independent campaigns. Screening failure and failure to optimize to nanomolar potency are different decisions with very different resource implications.
The text's precise Inh-31 value of 22.6 µM was not independently corroborated from the accessible primary abstract during this review. It should be checked in the full paper before reuse. This uncertainty is not the main basis for the score: the report appropriately does not treat Inh-31 as validated direct-binding chemistry.
Recommended use
Borrow the assay-qualification, donor-panel, tissue and decision-register structure. Replace the free-trough/whole-blood comparison with ChatGPT 6 Astra-high's matrix definitions. Reclassify dosing and competitive preferences as revisitable portfolio gates, and separate screening gates from optimization gates.
Most valuable contribution: a practical framework for managing evidence across discovery stages, provided its numerical rules are calibrated before use.
Claude-Opus5-extra
Operational-looking, internally inconsistent
TPP 10/20 | Strategy 12/20 | Substantial revision needed
Evidence anchors: PDF pp. 11-15, Tables 4-5 and Figure 5; pp. 18-19, Table 6; pp. 20-23, Tables 7-8.
What works
The report supplies an extensive assay inventory, paired probe counterscreens, orthogonal binding and site confirmation, early species testing, explicit failure criteria and a reasonably flexible beyond-rule-of-five pathway. It identifies the induced pocket as offset from the direct receptor footprint and gives useful peptide-derived anchor ideas.
Its distinction between downstream inhibitors and direct extracellular blockade is clear. The early species-coverage warning and emphasis on resynthesized material are good operational choices. Schematic figures are labeled as schematics; the lack of coordinate renderings is not itself a scientific defect.
What needs repair
The claimed TPP coherence is not demonstrated. Page 20 says the numbers are internally consistent with at most 100 mg QD, but Table 7 permits 300 mg QD and whole-blood IC50 up to 2 µM. It supplies no dose model connecting that potency to the allowed exposure. The dose-row rationale says the profile is below the approved peptide's 200 mg, which is false for the stated 300 mg minimum.
Coverage targets do not line up. Table 7 asks for free skin trough at cellular IC50, alongside 80% systemic proximal inhibition and an ideal IC90 coverage margin. Skin IC50 coverage does not establish 80-90% local suppression. The relevant matrix, tissue potency and exposure duration must be harmonized rather than treated as separate rows.
Mechanistic certainty exceeds the evidence. Pages 11 and 15 say receptor binding avoids competition with circulating cytokine. A competitive receptor antagonist still competes with cytokine; target choice does not remove ligand-concentration dependence. Page 12 also treats probe displacement as proof of access to the open pocket, although other binding modes, conformational effects and assay artifacts can produce displacement.
The tiered cascade lacks as explicit a native cytokine-receptor displacement requirement as the leading reports. Its live-cell receptor occupancy readout also needs adaptation for a cytokine-directed compound. The statement that cellular assays protect against cell-impermeable matter confuses target-cell entry with oral absorption for an extracellular target.
The “no public co-crystal of any ligand with IL23R” claim on p. 11 is too broad: the receptor-VHH complex 9MFG was released in 2025. [P7] Page 22's insistence that every toxicity finding must be attributable to chemistry overlooks on-target immune toxicity. A clinical PASI projection and a non-inferiority claim cannot be established from the proposed preclinical profile alone.
Recommended use
Reuse selected assay controls and exploratory chemistry ideas. Rebuild the TPP around one product objective and an explicit exposure model, add a direct native PPI gate, and remove categorical claims about competition, automatic selectivity and safety. Covalent proposals require a native, accessible residue and demonstrated chemistry before becoming a funded track.
Most valuable contribution: breadth of practical ideas; its weakness is the consistency and scientific qualification needed to choose among them.
PlayMolecule AI on Gemini-3.8-flash
Detailed methods do not rescue the premise
TPP 8/20 | Strategy 9/20 | Substantial revision needed
Evidence anchors: PDF pp. 6-12, Tables 2-4 and Figures 4-7; pp. 13-15, Table 5 and testing cascade; pp. 16-17, roadmap.
What works
PlayMolecule AI on Gemini-3.8-flash gives concrete laboratory options: a 2,000-3,000-fragment collection, SPR and NMR confirmation, off-DNA resynthesis, an IL-12 counterselection and a structured ADME checklist. The TPP has minimum/optimal columns and associated measurement methods. It is easier to turn individual rows into assay requests than the more clinical TPPs in PlayMolecule AI on GLM5.3-flash and PlayMolecule AI on Qwen3.8-28B, local, RTX 5090.
What needs repair
The central inhibition claim is unproven. Page 7 states that a small molecule occupying the induced pocket completely prevents receptor engagement. The primary paper instead attributes the peptide's blocking mechanism to its bulk at a site offset from the direct interface. A smaller binder may be silent. [P1] This is the first experiment to perform, not a conclusion on which to build the campaign.
The pharmacophore is presented too confidently. The four-subpocket scheme is a proposal, not an experimentally validated small-molecule pharmacophore. Claims that a particular amide geometry is critical for nanomolar affinity and that the estimated cavity accommodates rule-of-five ligands need actual non-peptide SAR and binding structures. A pocket-volume method and sensitivity to missing coordinates are not supplied.
There is also a reproducible contact error. Table 2 reports p19 Leu160 CD2 to IL23R Leu113 CD1 at 3.65 Å; the specified atoms in 5MZV are about 6.34 Å apart. Other checked distances in that table do reproduce. The lesson is to audit the atom-specific interaction map, not discard every structural observation. [P6]
The candidate checklist does not form a dose-feasibility model. There is no developed whole-blood potency target, unbound skin IC90 coverage requirement or human dose calculation. Vss and total skin partition are not substitutes for free drug in the extracellular target compartment. The species strategy is insufficiently qualified for the prescribed mouse efficacy package.
The cascade requires a positive thermal shift above 1.5°C, yet the supported fragments in the published supplement show negative shifts. A mandatory positive-shift rule could discard real binders; either shift direction still requires independent confirmation. [P1] Selectivity requirements also move between 500-fold, 100-fold and more than 1,000-fold without explaining stage or assay feasibility.
Calling seven-day imiquimod dermatitis a chronic translational model overstates what it establishes. The p. 16 assertion that meeting the TPP will ensure biologic-like clinical efficacy is not justified. Fourteen-day tolerability may inform nomination, but does not establish the chronic safety of the proposed product.
Recommended use
Keep the assay menu and selected property measurements. Rebuild the native-blockade hypothesis, structural interaction map, matrix-aware PK/PD model and species strategy before using the 18-month roadmap. Treat the timing as contingent on evidence, not a forecast.
Most valuable contribution: laboratory specificity; its central weakness is prescriptive detail attached to unvalidated mechanistic and translational assumptions.
PlayMolecule AI on GLM5.3-flash
Useful overview, insufficient candidate specification
TPP 7/20 | Strategy 10/20 | Substantial revision needed
Evidence anchors: PDF pp. 7-10, Tables 2-3 and Figure 6; p. 11, Table 4; pp. 13-14, staged plan and Table 6.
What works
PlayMolecule AI on GLM5.3-flash assembles relevant structures and useful analogies from other cytokine PPI programs. It recognizes induced-pocket discovery, combines fragment work with structural follow-up, includes mandatory IL-12 counterscreening, and gives a staged plan with potency/selectivity gates. It also acknowledges the need for unbound skin exposure and the cytoplasmic location of the protective R381Q variant.
Its overview of alternative modalities can help a portfolio discussion. That breadth should not be confused with a sufficiently specified non-peptide campaign.
What needs repair
The main modality drifts. The executive framing emphasizes the small-molecule opportunity, but pp. 9-10 explicitly recommend leading with a peptide/macrocycle program and running the small-molecule track in parallel. Macrocycle chemistry can remain non-peptide, but the report's display-derived peptide fallback and icotrokinra-like lead program do not fulfill the requested small-molecule objective. Peptides are valuable controls and templates; promoting them is a separate portfolio decision.
The TPP is predominantly a future clinical profile. Table 4 includes clinical PASI, onset, maintenance, infection rates, bioavailability, half-life and a 10-100 mg QD dose, but omits a coherent biochemical/cellular/whole-blood potency and target-site coverage specification. The later sub-100 nM and greater-than-100-fold IL-12 gates help, but do not supply assay matrix, maximal inhibition, human donor coverage or the dose bridge.
The competitive logic is inconsistent. The narrative says the program must clearly exceed approximately 55% PASI90, while the TPP minimum is 50% and its stretch range 55-60%. Calling 50-55% “guselkumab-like” is inconsistent with the report's own approximately 73% guselkumab benchmark. The 240 mg icotrokinra dosing reference also needs replacement by the verified 200 mg QD regimen. [P3]
Screening-to-mechanism decisions need sharper definition. The report uses the induced-pocket overlap illustration to imply blockade too readily. The specified DEL selection with receptor co-presentation needs a clear competition/counterselection design so that it does not preferentially retain binders to the assembled complex. Such binders might stabilize the interaction rather than inhibit it.
The computational workflow puts FEP prominently before a validated binding pose and congeneric SAR series are established. The first go/no-go criterion is a “druggable pocket,” but a predicted pocket score cannot settle experimental ligandability. Screening-scale and optimization-stage failures need separate criteria. Species matching should precede the prescribed animal models.
Recommended use
Use this report for target and modality orientation. Rewrite the campaign around a primary non-peptide p19 program, preserve peptides as tools or separately approved fallbacks, and replace Table 4 with distinct product and candidate profiles. Add explicit native PPI, human-cell and matrix-adjusted exposure gates.
Most valuable contribution: strategic breadth; it leaves too many implementation decisions unresolved to serve as the campaign specification.
PlayMolecule AI on Qwen3.8-28B, local, RTX 5090
Consequential interpretation errors
TPP 5/20 | Strategy 6/20 | Rebuild before operational use
Evidence anchors: PDF pp. 3-5, structural contacts and Table 2; pp. 11-13, pharmacology and strategy; pp. 16-18, Table 5 and risk section.
What works
PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 provides a broad structural and clinical survey, distinguishes oral peptides from small molecules, acknowledges the difficulty of the PPI, and includes species and human-cell work. It describes reporter-only and computational evidence as weaker than direct binding in parts of the report.
What requires correction first
The IFNγ interpretation is wrong and changes the assay strategy. Page 11 concludes that icotrokinra lacks selectivity over IFNγ because it suppresses IFNγ production at picomolar concentrations. The primary study measured IFNγ produced after IL-23 stimulation: suppression is the intended downstream effect. It is not evidence that the compound directly blocks the IFNγ pathway. The same paragraph also misidentifies the 18.4 pM NK-cell result as a hidradenitis-suppurativa cell line. This error is propagated into the TPP selectivity rationale. [P4]
The contact map contains incorrect design inputs. Table 2 assigns a 2.3 Å hydrogen bond to the W156 indole nitrogen and D118, with chemically reversed donor/acceptor language. In 5MZV, the indole NE1 is approximately 4.22 Å from D118 OD2 and 5.62 Å from OD1. The listed K164-G24 contact is 1.7 Å; the specified NZ-to-backbone-O distance is approximately 2.62 Å. These are not minor differences for a pharmacophore design. [P6]
A failed trial is assigned to the wrong asset. Pages 12 and 18 attribute NCT04102111 to icotrokinra/JNJ-77242113 and infer IBD futility risk from it. The registry intervention is JNJ-67864238, a different compound. Class attrition may still inform strategy, but it is not an icotrokinra failure. [P8]
The “preclinical” TPP is largely clinical. Week-16 clearance, week-52 durability and serious infections per patient-year are future human outcomes. Table 5 does not translate these into measured whole-blood potency, unbound trough coverage, a dose limit, property tradeoffs or feasible early discovery gates. A half-life alone does not establish once-daily efficacy. The “no exposure-safety” target is also not a meaningful general requirement for a new small molecule.
The report is inconsistent about pocket identity and priority: its primary loop-mimetic proposal is receptor-directed, while its conclusion calls the program p19-centric. Missing 8UUI coordinates are treated too readily as a resolved pocket-opening mechanism. Approval status is left unverified despite the report's 2026 context; this is less important than the mechanistic errors but weakens the product benchmark.
Recommended use
Retain the bibliography as a source-discovery aid and selectively recheck structural figures. Rebuild the pharmacology interpretation, atom-level design map and TPP from primary sources before directing experiments. The other reports offer better starting material for that rewrite.
Most consequential weakness: it can cause a team to reject desired on-target pharmacology as an off-target liability.
Cross-report findings that change decisions
This table separates independently checked corrections from judgments about what a discovery program should require.
| Finding | Evidence and affected reports | Operational consequence |
|---|---|---|
| Pocket binding is not established small-molecule antagonism | The 2025 paper describes an offset pocket and steric blockade by peptide bulk. PlayMolecule AI on Astra-high, ChatGPT 6 Astra-high and ChatGPT 5.6-Sol-ultra handle this well; PlayMolecule AI on Gemini-3.8-flash asserts blockade; PlayMolecule AI on GLM5.3-flash and Claude-Opus5-extra overstate parts of the inference. [P1] | Measure native cytokine-receptor inhibition separately from peptide displacement and direct binding. |
| Four highlighted fragment binders, not five | Supplement Table S2 gives hits 1/2/4/5 at 200/434/180/289 µM; hit 3 has no binding support. The main text gives a different range. PlayMolecule AI on Astra-high explicitly preserves this discrepancy; ChatGPT 6 Astra-high and ChatGPT 5.6-Sol-ultra repeat the broad range; PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 describes five confirmed hits. [P1] | Reproduce identities and affinities before starting an optimization series. |
| Thermal shift direction is not a universal gate | Supported fragments show negative DSF shifts in Table S2. PlayMolecule AI on Gemini-3.8-flash requires positive stabilization in its cascade. [P1] | Use thermal shifts as supportive observations, with independent binding and protein-integrity controls. |
| Human hotspot behavior is distributed | Human W156A alone is not a universal signaling knockout; combined mutation and experimental context matter. PlayMolecule AI on Astra-high and ChatGPT 5.6-Sol-ultra give the strongest qualification. [P5] | Avoid designing or interpreting mutants as if one human residue determines all antagonism. |
| IFNγ output is not IFNγ pathway selectivity | The peptide study's low-pM effect is on IL-23-induced IFNγ production. PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 interprets it as an off-target limitation. [P4] | Separate stimulation pathway, measured output and direct off-target signaling assays. |
| Some quoted atomic contacts do not reproduce | PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 W156/D118 and K164/G24 distances; PlayMolecule AI on Gemini-3.8-flash Leu160/Leu113 distance differ from the specified 5MZV atoms. [P6] | Recompute atom-specific contacts before commissioning hotspot-mimetic chemistry. |
| Direct oral competition already exists | FDA records March 17, 2026 approval. The label specifies 200 mg QD and fasting instructions. ChatGPT 6 Astra-high and ChatGPT 5.6-Sol-ultra use this most concretely; PlayMolecule AI on Astra-high is transparent about access limits. [P2, P3] | Require a plausible additional product advantage, but do not make any one convenience metric an unexplained universal kill rule. |
| Target-side selectivity is a hypothesis for the compound | p19 specificity supports IL-12 sparing, but a molecule may also bind other proteins. PlayMolecule AI on GLM5.3-flash, PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 and Claude-Opus5-extra use overly automatic language. | Test functional selectivity at efficacious exposure, not just binding-site identity. |
Stress-testing the TPP numbers
The main pharmacological distinction is between nominal concentration in an assay, free concentration in plasma, and free concentration in the skin compartment containing the target. They cannot be substituted for one another without an experimentally supported relationship.
A useful calculation, with explicit assumptions
For an illustrative full-inhibition curve with Hill slope 1, IC90 = 9 × IC50. If the cellular potency has been expressed as unbound concentration, a provisional relationship is:
Cmin,total,plasma = IC90,u / (fu,plasma × Kp,uu,skin)
Here fu is the unbound plasma fraction and Kp,uu is the unbound skin-to-plasma ratio. Using ChatGPT 6 Astra-high's assumptions of MW 500, fu 0.10, Kp,uu 0.50, F 0.50, human clearance 1 L/h and average total concentration twice the trough gives:
| Unbound cellular IC50 | Unbound IC90 | Total plasma trough | Approximate daily dose |
|---|---|---|---|
| 3 nM | 27 nM | 0.54 µM | 26 mg |
| 10 nM | 90 nM | 1.80 µM | 86 mg |
| 30 nM | 270 nM | 5.40 µM | 259 mg |
Dose is calculated as clearance × average concentration × 24 hours / bioavailability, with the molar-to-mass conversion. These are scenarios, not measured compound properties. A threefold coverage margin triples these doses; changing clearance, binding or distribution can change the conclusion substantially. A mechanistic PK/PD model may justify a different coverage rule when dissociation, receptor turnover or downstream persistence matters.
What this exposes in the reports
ChatGPT 6 Astra-high: the calculation is a major strength. Clarify the 80% maximum-effect floor versus an absolute IC90 requirement, and replace the assumed average/trough relationship with an actual PK model when data exist.
PlayMolecule AI on Astra-high: the right exposure framework is present, but the dose-feasibility sensitivity should be added. Its exposure-based counterpathway gate is stronger than a selectivity ratio alone.
ChatGPT 5.6-Sol-ultra: define whether the whole-blood IC90 is nominal or unbound before comparing it with free trough. A nominal whole-blood assay already incorporates its matrix effects; an additional unqualified binding correction can count those effects twice.
Claude-Opus5-extra: the 2 µM whole-blood IC50 ceiling, 100/300 mg dose claims and free-skin-IC50 gate are not a demonstrated coherent package. This does not prove that every such compound is infeasible; it means the report has not supplied the needed calculation.
PlayMolecule AI on Gemini-3.8-flash, PlayMolecule AI on GLM5.3-flash and PlayMolecule AI on Qwen3.8-28B, local, RTX 5090: bioavailability, half-life, Vss, total skin partition or an affinity threshold cannot independently establish the intended regimen. Each needs the same matrix-to-exposure-to-dose bridge.
Selectivity must be checked at the chosen dose
A 30-fold IC50 ratio is not automatically clean at an exposure selected to cover target IC90. In a deliberately simplified same-matrix, unit-Hill example, 3 × target IC90 equals 27 × target IC50. An off-target IC50 at 30 × target IC50 would then give about 47% off-target inhibition at that concentration. Different curve shapes and peak exposure alter the result. The point is to test the actual counterpathways at projected active exposures, not to declare a fixed ratio universally safe.
A starting package assembled from the best sections
The reports justify a gated feasibility campaign, not a claim that a non-peptide development candidate is already attainable. The following is my proposed synthesis for using them, rather than a new promise of performance or timing.
First decision: can reproducible matter block the native interaction?
Use PlayMolecule AI on Astra-high's compound-identity and supplement audit to define the initial replication set. Qualify intact human p19:p40 and appropriate receptor constructs with known peptide and antibody controls. Document chain/sequence mapping, protein activity, reagent concentrations, assay quality and soluble compound limits before screening.
Confirm binding using independent principles and fresh material. Include the distal probe, optical and colloidal controls, alternate immobilization where useful, chemical purity and mixture resolution. Weak fragments can initially be accepted as binding seeds; do not impose candidate-level nanomolar thresholds at this stage. Lack of saturated binding at the tested concentration should be reported as a limit, not a precise affinity.
Test native IL-23-IL23R disruption separately. A weak binding seed without measurable antagonism is not yet a validated antagonist; it may receive a bounded chemistry experiment to test growth toward the blocking volume. Stop treating the series as a lead if improved binding repeatedly fails to improve native displacement and selective human-cell activity.
Second decision: is the mechanism credible and optimizable?
Use the direct-binding, PPI and primary-cell evidence chain from ChatGPT 6 Astra-high and PlayMolecule AI on Astra-high. Add ChatGPT 5.6-Sol-ultra's assay-qualification and donor framework. Define the stimulus and readout for every cytokine assay; include IL-12, IL-6, type-I IFN, viability and direct JAK/TYK testing as appropriate. Define maximum inhibition and IC90 consistently.
Map the binding site with a co-structure where feasible, or convergent solution mapping, informative mutants and SAR. Mutants must preserve folding and interpretable baseline function. Peptide competition alone is insufficient to establish either an identical site or an identical blocking mechanism.
Make the induced p19 pocket the primary route. Keep a smaller, explicitly non-peptide receptor-side route if resources support it. Do not let peptide optimization, generic downstream inhibition, covalent labeling of engineered residues or binding to the preassembled complex silently replace the original objective.
Third decision: can the chemistry support a differentiated oral product?
Adopt ChatGPT 6 Astra-high's matrix definitions and sensitivity model; add measured binding, clearance, solubility, oral exposure and target-compartment distribution as the series matures. Use PlayMolecule AI on Astra-high's exposure-based functional selectivity gate. Choose efficacy and safety species only after candidate-specific cross-reactivity is known, and require human-cell/tissue pharmacology before nomination.
Use ChatGPT 5.6-Sol-ultra's decision register to record evidence, uncertainty and the next experiment. Separate failures of assay validity, chemical tractability, pharmacology, dose feasibility and product differentiation. Make commercial constraints explicit sponsor choices with tradeoffs, rather than universal scientific truths.
Practical selection: choose PlayMolecule AI on Astra-high if only one report can guide the first experimental package. Pair it with ChatGPT 6 Astra-high for a stronger TPP and PK/PD specification. Add selected ChatGPT 5.6-Sol-ultra execution tables after revision. Use the other reports as sources of hypotheses to check, rather than as instructions for spending the campaign budget.
Evaluation methods and original prompt
Exact report-generation prompt
The user supplied the following prompt as the common task used to generate the seven reports:
"Create an informed, validated report with text, tables, and figures based on publicly available information about this target, useful for drug discovery. Include structural information and visualizations, possible discovery strategies with small molecules, and a target product profile (TPP) of a small-molecule inhibitor therapeutic for the IL23R-p19 interaction for psoriasis at the preclinical level, and possible structural strategies to target it. "
What determines the ranking
An operational TPP links a desired product to measurable candidate criteria and a feasible exposure profile. A long list of potency, half-life, solubility and safety numbers is insufficient unless those values can be satisfied together.
An operational small-molecule strategy makes the central hypothesis falsifiable. For this target, the crucial chain is chemical identity, direct binding, native receptor displacement, selective human-cell inhibition, and exposure-linked skin pharmacology. A peptide-displacement signal or a modeled pocket does not establish the whole chain.
All 135 PDF pages were text-extracted and reviewed for relevant content; 34 decisive pages containing TPPs, strategies and figures were also rendered for visual inspection. Key claims were checked against primary papers, a supplement, public structure coordinates, regulatory records and a trial record. Instructions or recommendations inside the PDFs were treated as material to evaluate, not instructions to execute.
Independent calculations and boundaries
Atom checks used model 1, author chain identifiers B for p19 and C for IL23R in 5MZV, named atoms and Euclidean distances. Alternate atoms were selected by highest occupancy; the checks exclude hydrogen atoms and symmetry copies. Values are rounded to 0.01 Å for comparison, not as claims of experimental precision. Selected 8UUI contacts were also checked using author chain B for p19 and D for peptide. Electron-density maps were not re-refined, and no comprehensive pocket-volume, free-energy or ligandability calculation was performed.
Dose arithmetic and the selectivity example were independently recomputed from their stated assumptions. Scientific and commercial judgments are identified as such. No compounds were purchased, synthesized or tested; no external messages or laboratory instructions were sent. The comparison was completed against information accessible on 9 September 2026.
Sources and report locators
Supplied reports
Reports are identified by the system names in the supplied models.md mapping. All locators refer to the supplied PDF page index, beginning with the first page as page 1.
| Report (system) | Pages | Principal evaluation anchors |
|---|---|---|
| PlayMolecule AI on Astra-high | 14 | Strategy pp. 5-8; TPP pp. 9-10; gates p. 11. |
| ChatGPT 6 Astra-high | 16 | Strategy pp. 5-8; TPP pp. 9-10; dose p. 11; gates p. 12. |
| ChatGPT 5.6-Sol-ultra | 22 | Strategy pp. 12-16; TPP pp. 17-18; gates pp. 19-20. |
| Claude-Opus5-extra | 26 | Strategy pp. 11-15; assays pp. 18-19; TPP pp. 20-22; risks p. 23. |
| PlayMolecule AI on Gemini-3.8-flash | 18 | Structural claims pp. 6-12; TPP pp. 13-14; cascade pp. 15-17. |
| PlayMolecule AI on GLM5.3-flash | 15 | Strategy pp. 9-10; TPP p. 11; implementation pp. 13-14. |
| PlayMolecule AI on Qwen3.8-28B, local, RTX 5090 | 24 | Contacts pp. 3-5; pharmacology pp. 11-13; TPP pp. 16-18. |
Primary checks
P1. Lecomte et al. (2025), Identification of an Induced Orthosteric Pocket in IL-23, and supporting file cb5c00181_si_001.pdf. Main text and supplement retrieved through Europe PMC; Table S2 is on supplement PDF p. 20. Article and supplement access | Publisher DOI.
P2. FDA, Drug Trials Snapshots: ICOTYDE. Original approval date and trial-specific efficacy populations. FDA snapshot.
P3. ICOTYDE prescribing information, DailyMed. Dose, administration and safety context. Current label.
P4. Fourie et al. (2024), JNJ-77242113, a highly potent, selective peptide targeting the IL-23 receptor... Stimulus/readout interpretation and peptide pharmacology. Primary study.
P5. Deciphering site 3 interactions of interleukin 12 and interleukin 23 with their cognate murine and human receptors (2020). Human/mouse mutagenesis limitations. Primary study.
P6. RCSB PDB 5MZV and 8UUI. Public mmCIF coordinates downloaded for targeted contact checks.
P7. RCSB PDB 9MFG. Receptor-VHH co-crystal, released May 21, 2025; associated Ota et al. primary publication.
P8. ClinicalTrials.gov NCT04102111. Intervention identity and termination reason checked through the public study API.