Tran Thai Bao , Le Thi Thuy My , Mai Ngoc Qui , Pham Thi Bich Thao * and Nguyen Thanh Tien

* Corresponding author (ptbthao@ctu.edu.vn)

Abstract

Missing values and anomalous measurements are common in experimental materials datasets and can compromise both statistical reliability and the preservation of underlying physical relationships. While numerous data recovery methods have been developed, their ability to maintain the physical consistency of experimental data remains insufficiently explored. This study proposes a large-language-model-assisted workflow for processing zeolite heat-capacity data, in which the model supports data inspection, anomaly detection, code generation, and result visualization. These findings demonstrate that linear interpolation is more suitable for recovering zeolite heat-capacity data when both data completeness and physical consistency are required. Furthermore, the proposed workflow highlights the potential of large language models as effective research assistants for semi-automated processing, analysis, and visualization of experimental materials data while maintaining human oversight in scientific interpretation.

Keywords: Heat capacity, linear interpolation, missing data, zeolite; agent

Tóm tắt

Dữ liệu thực nghiệm về vật liệu thường chứa các giá trị thiếu và giá trị sai lệch, gây ảnh hưởng đến độ tin cậy của các phân tích thống kê và khả năng bảo toàn ý nghĩa vật lý của dữ liệu. Hiện nay, phần lớn các nghiên cứu tập trung vào khả năng khôi phục dữ liệu mà chưa đánh giá đầy đủ mức độ duy trì các mối quan hệ vật lý sau xử lý. Nghiên cứu này được thực hiện nhằm đề xuất một quy trình xử lý dữ liệu nhiệt dung zeolite dựa trên mô hình ngôn ngữ lớn trong các bước kiểm tra dữ liệu, phát hiện dữ liệu lỗi, hỗ trợ sinh mã xử lý và trực quan hóa kết quả. Kết quả nghiên cứu cho thấy nội suy là phương pháp phù hợp đối với dữ liệu nhiệt dung zeolite, đáp ứng đồng thời việc khôi phục dữ liệu và bảo toàn các đặc trưng vật lý của hệ. Kết quả cũng minh chứng tiềm năng của mô hình ngôn ngữ lớn như một tác nhân hỗ trợ hiệu quả trong quy trình xử lý và phân tích dữ liệu thực nghiệm vật liệu theo hướng bán tự động và có khả năng tái lập.

Từ khóa: Dữ liệu khuyết, dữ liệu sai lệch, nhiệt dung, tác nhân, zeolite

Article Details

References

Ball, P. (2006). Cold comfort. Nature Materials, 5(3), 174–174.
https://doi.org/10.1038/nmat1602

Bazgir, A., Madugula, R. C. P., & Zhang, Y. (2025, April). MatAgent: A human-in-the-loop multi-agent LLM framework for accelerating the material science discovery cycle [Paper presentation]. AI for Accelerated Materials Design Workshop, ICLR 2025, Singapore.
https://openreview.net/forum?id=2Nm6Ef4tZD

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In M. C. Elish, W. Isaac, & R. S. Zemel (Eds.), Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). Association for Computing Machinery.
https://doi.org/10.1145/3442188.3445922

Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624(7992), 570–578.
https://doi.org/10.1038/s41586-023-06792-0

Butler, K. T., Davies, D. W., Cartwright, H., Isayev, O., & Walsh, A. (2018). Machine learning for molecular and materials science. Nature, 559(7715), 547–555. https://doi.org/10.1038/s41586-018-0337-2

Curtarolo, S., Setyawan, W., Hart, G. L. W., Jahnatek, M., Chepulskii, R. V., Taylor, R. H., Wang, S., Xue, J., Yang, K., Levy, O., Mehl, M. J., Stokes, H. T., Demchenko, D. O., & Morgan, D. (2012). AFLOW: An automatic framework for high-throughput materials discovery. Computational Materials Science, 58, 218–226. https://doi.org/10.1016/j.commatsci.2012.02.005

Jablonka, K. M., Ai, Q., Al-Feghali, A., Badhwar, S., Bocarsly, J. D., Bran, A. M., Bringuier, S., Brinson, L. C., Choudhary, K., Circi, D., Cox, S., de Jong, W. A., Evans, M. L., Gastellu, N., Genzling, J., Gil, M. V., Gupta, A. K., Hong, Z., Imran, A., … Blaiszik, B. (2023). 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digital Discovery, 2(5), 1233–1250. https://doi.org/10.1039/D3DD00113J

Jablonka, K. M., Ongari, D., Moosavi, S. M., & Smit, B. (2020). Big-Data Science in Porous Materials: Materials Genomics and Machine Learning. Chemical Reviews, 120(16), 8066–8129. https://doi.org/10.1021/acs.chemrev.0c00004

Jain, A., Ong, S. P., Hautier, G., Chen, W., Richards, W. D., Dacek, S., Cholia, S., Gunter, D., Skinner, D., Ceder, G., & Persson, K. A. (2013). Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Materials, 1(1), 011002-(1-11). https://doi.org/10.1063/1.4812323

Bran, AM., Cox, S., Schilter, O., Baldassari, C., White, A. D., & Schwaller, P. (2024). Augmenting large language models with chemistry tools. Nature Machine Intelligence, 6(5), 525–535.
https://doi.org/10.1038/s42256-024-00832-8

Rajan, K. (2005). Materials informatics. Materials Today, 8(10), 38–45. https://doi.org/10.1016/S1369-7021(05)71123-8

Ramprasad, R., Batra, R., Pilania, G., Mannodi-Kanakkithodi, A., & Kim, C. (2017). Machine learning in materials informatics: recent applications and prospects. Npj Computational Materials, 3(1), 54. https://doi.org/10.1038/s41524-017-0056-5

Rousseeuw, P. J., & van Zomeren, B. C. (1990). Unmasking Multivariate Outliers and Leverage Points. Journal of the American Statistical Association, 85(411), 633–639. https://doi.org/10.1080/01621459.1990.10474920

Saal, J. E., Kirklin, S., Aykol, M., Meredig, B., & Wolverton, C. (2013). Materials Design and Discovery with High-Throughput Density Functional Theory: The Open Quantum Materials Database (OQMD). JOM, 65(11), 1501–1509.
https://doi.org/10.1007/s11837-013-0755-4

Schafer, J. L., & Graham, J. W. (2002). Missing data: Our view of the state of the art. Psychological Methods, 7(2), 147–177. https://doi.org/10.1037/1082-989X.7.2.147

Schmidt, J., Marques, M. R. G., Botti, S., & Marques, M. A. L. (2019). Recent advances and applications of machine learning in solid-state materials science. Npj Computational Materials, 5(1), 83.
https://doi.org/10.1038/s41524-019-0221-0

Stekhoven, D. J., & Bühlmann, P. (2012). MissForest—non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1), 112–118. https://doi.org/10.1093/bioinformatics/btr597

Ward, L., Agrawal, A., Choudhary, A., & Wolverton, C. (2016). A general-purpose machine learning framework for predicting properties of inorganic materials. Npj Computational Materials, 2(1), 16028. https://doi.org/10.1038/npjcompumats.2016.28

Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), 160018. https://doi.org/10.1038/sdata.2016.18

Xie, T., & Grossman, J. C. (2018). Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties. Physical Review Letters, 120(14), 145301. https://doi.org/10.1103/PhysRevLett.120.145301