Methodology
Research Pipeline — From Raw Data to Statistical Evidence
Data Collection
ecoinvent 3.11
25,412 records
25,412 records
→
Preprocessing
Feature engineering
NLP encoding
NLP encoding
→
Clustering
KMeans · Agglom.
DBSCAN · k=8
DBSCAN · k=8
→
Anomaly Detect.
Isolation Forest
Local Outlier Factor
Local Outlier Factor
→
LLM Benchmark
6 models · 2 task
types · n=250
types · n=250
→
Statistical Tests
Cochran's Q
McNemar · Bonferroni
McNemar · Bonferroni
→
Results
Validated findings
p < 0.001
p < 0.001
Objective 3
LLM Benchmarking — The Accuracy Illusion Exposed
THE ACCURACY COLLAPSE
// ALL MODELS NEAR-PERFECT · THEN WE RAISED THE DIFFICULTY //
3/6perfect structural completeness on ISO 14044 Goal & Scope
9/10Claude Haiku 4.5 comparative assertions — best across all models
p<0.001Cochran's Q confirms statistically significant differences
−65 ppGPT-4o-mini: maximum accuracy collapse discovered
Multi-Dimensional Model Comparison
Task 2 — ISO 14044 Goal & Scope Drafting
STRUCTURAL COMPLETENESS
COMPARATIVE ASSERTIONS (/ 10)
Claude Haiku 4.5 led Task 2 overall (9/10 assertions) despite ranking 3rd on Task 1b — distinct capability profiles across task types.
Explorer
What If? — Accuracy vs. Task Difficulty
Drag the slider to explore how accuracy degrades with difficulty
Rankings
Live Leaderboard — Watch Rankings Reshape
Click model name for full profile
Battle Mode
Head-to-Head — Pick Any Two Models & Fight
VS
Collapse Map
Easy vs Hard — The Accuracy Collapse Visualised
Bubble size = accuracy drop (larger = worse collapse). Points above the diagonal = consistent; below = collapsed.
Statistics
Pairwise Statistical Significance — McNemar Test
McNemar's test on pairwise model disagreements (hard task, n=50).
Green = significant difference (p<0.05) ·
Red = not significant. Hover each cell for exact p-value.
p < 0.001 — highly significant
p < 0.01 — significant
p < 0.05 — marginally significant
p ≥ 0.05 — not significant
PRISMA Clusters
Research Landscape — 8 Thematic Clusters from 209 Papers
Sentence-BERT + UMAP + HDBSCAN · Word size = paper count · Hover to explore
Objective 2
Unsupervised ML on 25,412 ecoinvent Records
Clustering Method Performance
DBSCAN best silhouette (0.792) via noise identification. K-Means and Agglomerative independently converged on k=8 clusters. Three methods triangulate a validated partition.
Anomaly Detection — Agreement Analysis
0Isolation
Forest
Forest
∩0shared
0Local
Outlier Factor
Outlier Factor
Jaccard Similarity = 0.001
Root cause: 23.2% exact-duplicate feature vectors corrupted LOF's density landscape. Confirmed high-impact outliers: aviation and land-use change processes.
Objective 1
Landscape of AI in LCA Research & Commercial Software
PRISMA Systematic Review — Paper Selection Flow
0
papers identified (database search)
▼ title & abstract screening
0
relevant to AI-in-LCA (PRISMA screen)
▼ full-text eligibility check
0
full-text papers included
▼ Sentence-BERT + UMAP + HDBSCAN
0
thematic clusters identified
AI in LCA research concentrates overwhelmingly on LCI stage. Goal & Scope drafting and interpretation tasks remain almost entirely uncharted territory.
8 Commercial AI-LCA Platforms (2026)
MakersiteAI gap-filling; automated BOM-to-database matching
SpheraPredictive matching to proprietary GaBi database
One Click LCAAI mapping of BIM/BOQ files to EPD datasets
MinviroData-driven parameterisation for geological variables
Muir AILLM-driven synthetic supply-chain deconstruction
CarbonCloudAI classification for agricultural supply chains
WatershedSpend-based emissions estimation & integration
TerrascopeAutomated Scope 3 calculation and reporting
Publications
Research Output
SSRN Preprint
Understanding the Role of Artificial Intelligence in Life Cycle Assessment
Coming Soon
A preprint of the full dissertation findings will be deposited to SSRN once grading is complete. Will include all three objectives, statistical results, and LLM benchmark data.
Full Dissertation Report
MSc Dissertation — University of Warwick, WMG · September 2026
Coming Soon
The complete dissertation document (approx. 20,000 words) covering PRISMA review, ecoinvent ML analysis, and LLM benchmarking with McNemar statistical tests.
Journal / Conference Paper
AI-Driven Goal & Scope Drafting in LCA: A Benchmarking Study of Large Language Models
Coming Soon
A condensed peer-review submission targeting the accuracy collapse finding and McNemar significance results across six frontier LLMs. Venue TBD post-result.
"
Central methodological principle · Raj Khatik (2026), §1.4 · Demonstrated across Objectives 2 and 3