Bin Cao and colleagues introduce Gan Jiang, a self-learning scientific agent for powder X-ray diffraction (PXRD). Their paper, A self-learning scientific agent for X-ray diffraction, was released as an arXiv preprint on 6 October 2026.
Gan Jiang connects candidate retrieval, multiphase analysis, physics-constrained modeling, and result verification into one workflow. It records failure diagnoses and effective operations in an external skill library, allowing validated improvements to carry forward to new samples while the language model’s parameters remain unchanged.
Using XRD patterns alone
on held-out samples after freezing
From complex phases to evolving lattices
From a diffraction pattern to a credible structure
A powder X-ray diffraction pattern encodes information about a material’s crystal structure. Peak positions constrain interplanar spacings; relative intensities relate to atomic arrangements and phase fractions; peak widths also reflect crystallite size, strain, and instrumental effects. The task is to turn these signals into testable structural judgments.
Powder diffraction compresses scattering from many crystallites into a one-dimensional curve. Overlapping peaks, coexisting phases, structural disorder, and incomplete measurements can make different models appear similarly plausible. Analysis often involves repeated candidate searches, inspection of unexplained peaks, refinement adjustments, and checks against physical and chemical constraints.
This study asks a further question: how can experience from one analysis become a method that the next analysis can use? Gan Jiang preserves tool calls, residual diagnoses, and procedural corrections, making the analysis procedure itself something that can be tested, revised, and reused.
Every step connected to evidence
Gan Jiang separates language-model planning, specialist computation, and scientific verification. The language model interprets the research objective and selects skills; diffraction engines perform retrieval, decomposition, and refinement; explicit criteria determine whether to accept a candidate, retry, or take another branch.
Workflow adapted from Figs. 2b and 3 of the paper. In automatic mode, multiphase analysis begins only after a single-phase hypothesis has undergone full verification and been rejected. A failed tool call alone does not establish that a sample is multiphase.
Four complementary analysis engines
XMatcher · Interpretable pattern matching
Retrieves candidate structures from observed diffraction peaks while retaining peak correspondences and residual signals, so researchers can inspect which observations support a candidate and which remain unexplained.
XQueryer-LW · Learned structure retrieval
When elemental information is supplied, this engine combines the pattern with elemental constraints to retrieve complementary candidates. This retrieval channel is not enabled without elemental input.
XDecomposer · Multiphase decomposition
Estimates component diffraction signals in mixed patterns. XMatcher-AutoMix also supports searches over candidate phase combinations. Each combination must still pass joint evidence checks.
PyWPEM · Physics-constrained modeling
Converts structural hypotheses into calculated diffraction patterns, using the full curve to constrain the model. The system also integrates FullProf and GSAS-II as refinement branches.
The candidate structure library defines the retrieval space: 100,315 MP500 structures and 2,007 experimentally sourced structures. This library serves retrieval, while the separate DeltaXRDbench pattern collection serves performance evaluation.
A critical detail: preserve the original experimental signal
Peak analysis and refinement use the full measured angular range and original intensity scale. XQueryer-LW receives a separately generated, normalized 3,500-point input spanning 10°–90°, with unmeasured regions treated only as padding. XRDinspector receives the full-range data separately. Retrieval preprocessing therefore does not overwrite the experimental evidence used to assess candidates.
During identification, XRDinspector checks observed-peak coverage, support for predicted peaks, evidence for each phase, and elemental consistency. During refinement, it compares observed and calculated curves and examines local residuals. Inputs, tool responses, and policy versions are retained so structural judgments can be traced to their execution history.
Passing verification indicates agreement between the model and pattern under the stated criteria; it does not establish a unique structural interpretation. Limited angular-offset trials also do not replace instrument calibration.
Turning analysis experience into executable skills
Gan Jiang stores long-term experience outside the model. A skill combines instructions, prompts, and supporting code to specify tool calls, parameter-release order, required diagnostics, and failure handling. Adjusting lattice parameters improves one sample’s fit; revising the parameter-release sequence for later samples improves the analysis method itself.
Validate the improvement, then test transfer
Starting from expert-optimized workflows, training cases provide execution traces and diffraction feedback. Old and new versions are compared in pairs on a development set. Only after a selected skill version is frozen is it evaluated on held-out samples. Held-out samples do not contribute to skill revision or version selection.
| Refinement branch | Initial score | Final score | Score gain |
|---|---|---|---|
| FullProf | 46.90 | 72.38 | +25.48 |
| GSAS-II | 39.01 | 60.02 | +21.01 |
| PyWPEM | 31.74 | 38.33 | +6.59 |
Source: Fig. 4 and the main text. Scores reflect diffraction-pattern fitting performance, with failures and timeouts counted as zero; they are not structure-identification accuracies. Allowed revisions differ across branches, so gains should be interpreted within each branch.
What changed in practice
Generic staged tuning was replaced with task-specific searches over peak widths and angular offsets. Explicit bounds, step-size reduction rules, and continuation criteria made the revised procedure more precisely executable.
New checks confirm that parameters reach the tool and that saved results match the final calculated curve. Following a training case in which adjusting several parameter groups at once worsened the fit, the skill also retained the practice of starting from defaults and adjusting one group at a time.
For tasks with an existing CIF structure, unnecessary structure-solution steps are skipped by default. Targeted handling of divergence and timeouts reduces mismatches between the execution path and the analysis objective.
These experiments show that controlled workflow revisions can improve analysis of samples excluded from version selection. The deployment architecture also includes regression checks, version control, monitoring, and rollback as a basis for continued learning in use. Long-term, open-ended self-learning remains a subject for further testing.
Testing identification on simulated and experimental data
The study introduces DeltaXRDbench, a benchmark containing 62,044 patterns: 12,044 single-phase patterns and 50,000 two- or three-phase mixtures constructed from source patterns. It spans simulated MP500 data, experimental mineral data from RRUFF, and experimental opXRD data, covering different origins and measurement complexities.
Evaluation matches submitted CIFs against reference crystal structures instead of relying solely on database identifiers. Single-phase tasks assess candidate rankings; multiphase tasks assess whether the complete, correct set of phases has been recovered.
Leading single-phase Top-1 accuracy with XRD alone
Top-1 accuracy is the fraction of samples for which the first candidate structure is correct. Under the methods and settings evaluated in the paper, Gan Jiang achieves the highest result on all three sources. The chart compares it with AutoXRD, the strongest comparator for this metric.
Single-phase Top-1 accuracy · XRD only
Gan Jiang and AutoXRD compared under the same input condition. Gains are percentage points (pp).
Adapted from Table 1. Single-phase sample counts are 10,000 for MP500, 1,164 for RRUFF, and 880 for opXRD. Gains are expressed in percentage points.
With reference composition information, Gan Jiang’s single-phase Top-1 accuracies rise to 99.90%, 92.16%, and 48.93%, respectively. The paper also reports Top-3, Top-5, and MRR@5, with Gan Jiang achieving the highest scores under the stated settings, extending the advantage beyond the first candidate.
Partial recovery and complete phase recovery
For samples containing multiple crystalline phases, identifying some components differs from recovering them all. Macro-F1 accounts for false positives and missed phases. Exact phase-set recovery requires every true phase to be found, with no extra phases. For a sample containing A, B, and C, finding only A and B can earn partial-recovery credit, but Exact remains zero.
| Data source | F1 · XRD only | F1 · + composition | Exact · XRD only | Exact · + composition |
|---|---|---|---|---|
| MP500 | 66.98% | 91.61% | 22.80% | 69.70% |
| RRUFF | 56.23% | 78.56% | 15.90% | 38.20% |
| opXRD | 27.04% | 38.21% | 3.10% | 8.20% |
Source: Table 2. Each evaluation draws 1,000 mixed samples per source; reported values are means over three fixed random seeds. Missing, invalid, or empty predictions receive zero identification credit.
Under both input conditions, Gan Jiang leads the evaluated methods in multiphase precision, recall, F1, and Exact. The absolute opXRD results also show that complete phase recovery remains difficult with incomplete angular coverage and complex experimental signals, identifying a concrete direction for improvement.
Benchmark mixtures are constructed by combining source signals and cannot represent every effect present in real mixtures. Composition-assisted protocols also differ across methods; comparisons should be read together with the settings specified in the paper.
Four materials problems, one analysis framework
From densely overlapping reflections to evolving lattices, Gan Jiang applies the same hypothesis–computation–verification framework to different materials problems, producing structural models and quantitative results that can be inspected further.
PbSO₄: constraining structure with the full pattern
Orthorhombic PbSO₄ (space group Pnma) contains 383 pairs of overlapping Bragg reflections in its pattern. Assigning isolated peaks can be ambiguous in such data. Whole-pattern modeling retains the combined constraints of peak positions, profiles, and intensity distributions.
The refined lattice parameters are approximately a = 8.4851 Å, b = 5.4016 Å, and c = 6.9643 Å. The candidate becomes a model that can be compared point by point with the observed curve. Lower residuals indicate closer pattern agreement under the respective definitions; structural judgments still require physical plausibility.
Ancient Egyptian cosmetics: resolving five minerals
For a synchrotron diffraction pattern of ancient Egyptian cosmetic powder, Gan Jiang separates and quantifies gypsum, phosgenite, cerussite, galena, and laurionite. The five-phase refinement yields mass fractions with a whole-pattern fit of Rwp = 13.079%.
Refined mass fractions sum to 100.00%. Data from the main text and Fig. 1b.
The substantial fraction of chlorine-bearing lead minerals is consistent with earlier evidence for artificial chemical processing, linking phase analysis to ancient materials preparation. These are mass fractions from structural refinement; peak-signal contributions estimated during candidate matching should be distinguished from mass or volume fractions.
Battery cathodes: from sequential scans to lattice trajectories
For an O3-type Li(Ni₀.₈Co₀.₁Mn₀.₁)O₂ cathode cycled over 3.0–4.6 V, Gan Jiang treats consecutive scans as a connected sequence. The verified structure from one frame initializes the next, while background and diffraction parameters are reassessed for the current measurement.
The recovered c-axis trajectory shows expansion followed by contraction at high voltage, consistent with measured peak shifts. Sequential analysis connects electrochemical state with lattice response, supporting the study of structural changes during operation.
Ru–Mn oxides: testing alternative Ru site occupancies
For a Ru–Mn oxide electrocatalyst prepared by cation exchange from orthorhombic Mn₂O₃, the question concerns atomic configurations: which Ru site occupations best explain the measurements? Gan Jiang screens candidate configurations in a 3 × 3 supercell and compares their whole-pattern refinements.
Within the tested configuration space, the result provides a diffraction-supported structural model for subsequent catalytic studies. It is a testable structural hypothesis; specific catalytic mechanisms still require complementary experiments and calculations.
Let analysis methods improve with the evidence
Gan Jiang brings physical analysis tools, agent workflows, and evaluation into a shared chain of verification. Tools produce calculated results, measurements test structural hypotheses, and held-out samples test workflow revisions. Both successful and failed analyses can contribute skills for future work.
What researchers provide
- PXRD patterns and the analysis objective
- Optional elements, synthesis context, and instrument conditions
- Existing CIF crystal-structure files
- Measurement range, wavelength, and background information
What the analysis returns
- Candidate structures or refined CIF files
- Observed and calculated patterns, plus refinement parameters
- Peak evidence, residuals, and execution diagnostics
- Traceable analysis reports and process records
The system supports phase identification, identification followed by refinement, and direct refinement of supplied structures. Researchers can use the outputs to revisit unexplained peaks, inspect structural parameters, and decide whether further measurements are needed.
Next: connect analysis feedback to experimental acquisition
When measurements lack features needed to distinguish candidates, further parameter tuning may not resolve the ambiguity. The paper proposes an agent that identifies missing or uncertain diffraction information, guides additional acquisition by the diffractometer, and updates the structural interpretation. This analysis–experiment loop still requires prospective experimental validation.
Use physical computation to support structural judgments, controlled tests to validate workflow improvements, and experience to inform the next analysis.
Authors Bin Cao, Huichi Zhou, Runyu Yang, Jingsong Li, Shuchen Sun, Yan Song, Hanyu Gao, Zhongwei Yu, Tong-Yi Zhang, and Jun Wang.
Co-first authors Bin Cao and Huichi Zhou
Corresponding authors Bin Cao, Tong-Yi Zhang, and Jun Wang
Version arXiv:2610.07862v1 · 6 October 2026
Contact bincao@hkust-gz.edu.cn
Bin Cao’s website bincao.work
All results and case studies refer to the paper version above. Figures 1–4 reproduce original paper figures; workflow summaries and data charts are adapted from the paper.