Procedure for processing and recovering experimental data of specific heat zeolite with the support of a large language model
Abstract
Missing values and anomalous measurements are common in experimental materials datasets and can compromise both statistical reliability and the preservation of underlying physical relationships. While numerous data recovery methods have been developed, their ability to maintain the physical consistency of experimental data remains insufficiently explored. This study proposes a large-language-model-assisted workflow for processing zeolite heat-capacity data, in which the model supports data inspection, anomaly detection, code generation, and result visualization. These findings demonstrate that linear interpolation is more suitable for recovering zeolite heat-capacity data when both data completeness and physical consistency are required. Furthermore, the proposed workflow highlights the potential of large language models as effective research assistants for semi-automated processing, analysis, and visualization of experimental materials data while maintaining human oversight in scientific interpretation.
Tóm tắt
Dữ liệu thực nghiệm về vật liệu thường chứa các giá trị thiếu và giá trị sai lệch, gây ảnh hưởng đến độ tin cậy của các phân tích thống kê và khả năng bảo toàn ý nghĩa vật lý của dữ liệu. Hiện nay, phần lớn các nghiên cứu tập trung vào khả năng khôi phục dữ liệu mà chưa đánh giá đầy đủ mức độ duy trì các mối quan hệ vật lý sau xử lý. Nghiên cứu này được thực hiện nhằm đề xuất một quy trình xử lý dữ liệu nhiệt dung zeolite dựa trên mô hình ngôn ngữ lớn trong các bước kiểm tra dữ liệu, phát hiện dữ liệu lỗi, hỗ trợ sinh mã xử lý và trực quan hóa kết quả. Kết quả nghiên cứu cho thấy nội suy là phương pháp phù hợp đối với dữ liệu nhiệt dung zeolite, đáp ứng đồng thời việc khôi phục dữ liệu và bảo toàn các đặc trưng vật lý của hệ. Kết quả cũng minh chứng tiềm năng của mô hình ngôn ngữ lớn như một tác nhân hỗ trợ hiệu quả trong quy trình xử lý và phân tích dữ liệu thực nghiệm vật liệu theo hướng bán tự động và có khả năng tái lập.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
References
Ball, P. (2006). Cold comfort. Nature Materials, 5(3), 174–174.
https://doi.org/10.1038/nmat1602
Bazgir, A., Madugula, R. C. P., & Zhang, Y. (2025, April). MatAgent: A human-in-the-loop multi-agent LLM framework for accelerating the material science discovery cycle [Paper presentation]. AI for Accelerated Materials Design Workshop, ICLR 2025, Singapore.
https://openreview.net/forum?id=2Nm6Ef4tZD
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In M. C. Elish, W. Isaac, & R. S. Zemel (Eds.), Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). Association for Computing Machinery.
https://doi.org/10.1145/3442188.3445922
Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624(7992), 570–578.
https://doi.org/10.1038/s41586-023-06792-0
Butler, K. T., Davies, D. W., Cartwright, H., Isayev, O., & Walsh, A. (2018). Machine learning for molecular and materials science. Nature, 559(7715), 547–555. https://doi.org/10.1038/s41586-018-0337-2
Curtarolo, S., Setyawan, W., Hart, G. L. W., Jahnatek, M., Chepulskii, R. V., Taylor, R. H., Wang, S., Xue, J., Yang, K., Levy, O., Mehl, M. J., Stokes, H. T., Demchenko, D. O., & Morgan, D. (2012). AFLOW: An automatic framework for high-throughput materials discovery. Computational Materials Science, 58, 218–226. https://doi.org/10.1016/j.commatsci.2012.02.005
Jablonka, K. M., Ai, Q., Al-Feghali, A., Badhwar, S., Bocarsly, J. D., Bran, A. M., Bringuier, S., Brinson, L. C., Choudhary, K., Circi, D., Cox, S., de Jong, W. A., Evans, M. L., Gastellu, N., Genzling, J., Gil, M. V., Gupta, A. K., Hong, Z., Imran, A., … Blaiszik, B. (2023). 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digital Discovery, 2(5), 1233–1250. https://doi.org/10.1039/D3DD00113J
Jablonka, K. M., Ongari, D., Moosavi, S. M., & Smit, B. (2020). Big-Data Science in Porous Materials: Materials Genomics and Machine Learning. Chemical Reviews, 120(16), 8066–8129. https://doi.org/10.1021/acs.chemrev.0c00004
Jain, A., Ong, S. P., Hautier, G., Chen, W., Richards, W. D., Dacek, S., Cholia, S., Gunter, D., Skinner, D., Ceder, G., & Persson, K. A. (2013). Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Materials, 1(1), 011002-(1-11). https://doi.org/10.1063/1.4812323
Bran, AM., Cox, S., Schilter, O., Baldassari, C., White, A. D., & Schwaller, P. (2024). Augmenting large language models with chemistry tools. Nature Machine Intelligence, 6(5), 525–535.
https://doi.org/10.1038/s42256-024-00832-8
Rajan, K. (2005). Materials informatics. Materials Today, 8(10), 38–45. https://doi.org/10.1016/S1369-7021(05)71123-8
Ramprasad, R., Batra, R., Pilania, G., Mannodi-Kanakkithodi, A., & Kim, C. (2017). Machine learning in materials informatics: recent applications and prospects. Npj Computational Materials, 3(1), 54. https://doi.org/10.1038/s41524-017-0056-5
Rousseeuw, P. J., & van Zomeren, B. C. (1990). Unmasking Multivariate Outliers and Leverage Points. Journal of the American Statistical Association, 85(411), 633–639. https://doi.org/10.1080/01621459.1990.10474920
Saal, J. E., Kirklin, S., Aykol, M., Meredig, B., & Wolverton, C. (2013). Materials Design and Discovery with High-Throughput Density Functional Theory: The Open Quantum Materials Database (OQMD). JOM, 65(11), 1501–1509.
https://doi.org/10.1007/s11837-013-0755-4
Schafer, J. L., & Graham, J. W. (2002). Missing data: Our view of the state of the art. Psychological Methods, 7(2), 147–177. https://doi.org/10.1037/1082-989X.7.2.147
Schmidt, J., Marques, M. R. G., Botti, S., & Marques, M. A. L. (2019). Recent advances and applications of machine learning in solid-state materials science. Npj Computational Materials, 5(1), 83.
https://doi.org/10.1038/s41524-019-0221-0
Stekhoven, D. J., & Bühlmann, P. (2012). MissForest—non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1), 112–118. https://doi.org/10.1093/bioinformatics/btr597
Ward, L., Agrawal, A., Choudhary, A., & Wolverton, C. (2016). A general-purpose machine learning framework for predicting properties of inorganic materials. Npj Computational Materials, 2(1), 16028. https://doi.org/10.1038/npjcompumats.2016.28
Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), 160018. https://doi.org/10.1038/sdata.2016.18
Xie, T., & Grossman, J. C. (2018). Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties. Physical Review Letters, 120(14), 145301. https://doi.org/10.1103/PhysRevLett.120.145301