Tech Logic / Intelligence Frontier

Exploring the Role of Large Language Models in the Scientific Method: From Hypothesis to Discovery

Large Language Models (LLMs) are transforming various stages of scientific research, including experimental design, data analysis, and hypothesis generation. This paper reviews the current application of LLMs in the scientific method, analyzes the gap between their roles as technical tools and creative engines, and points out the key steps required for deeper integration. Although LLMs show great potential in accelerating scientific discovery, they still face limitations in fundamental science (e.g., the discovery of new principles or laws). In the future, combining data-driven techniques with symbolic systems may give rise to hybrid engines that drive entirely new research directions.

TSO brief

  • Large Language Models (LLMs) are transforming various stages of scientific research, including experimental design, data analysis, and hypothesis generation. This paper reviews the current application of LLMs in the scientific method, analyzes the gap between their roles as technical tools and creative engines, and points out the key steps required for deeper integration. Although LLMs show great potential in accelerating scientific discovery, they still face limitations in fundamental science (e.g., the discovery of new principles or laws). In the future, combining data-driven techniques with symbolic systems may give rise to hybrid engines that drive entirely new research directions.
  • Tech Logic · Intelligence Frontier
  • Jul 26, 2026
TSO noteEach article is checked against independent reporting. The original source links are listed with the analysis so readers can inspect the evidence directly.

Source transparency

Original reporting sources

  1. 探索大型语言模型在科学方法中的作用:从假设到发现www.nature.com

Introduction

As recent Nobel Prizes recognize AI's contributions to science, large language models (LLMs) are transforming scientific research by boosting productivity and reshaping scientific methods. Today, LLMs are already involved in experimental design, data analysis, and workflows, with particularly notable impacts in chemistry and biology.

Recent advances in artificial intelligence have transformed many aspects of society, the global economy, and academic research practices. Generative AI and LLMs bring unprecedented opportunities to transform scientific practice, drive scientific development, and accelerate technological innovation. The Nobel Prizes in Physics and Chemistry have been awarded to several AI leaders, recognizing their contributions to AI and cutting-edge models such as LLMs. This signals that LLMs will transform or advance scientific research by enhancing productivity and supporting all stages of the scientific method. The application of AI in science is burgeoning, spanning numerous scientific fields and influencing different parts of the scientific process.

Despite the great potential of LLMs in hypothesis generation and data synthesis, AI and LLMs still face challenges in fundamental science and scientific discovery. Therefore, our perspective is that, so far, AI's impact on fundamental science—defined here as the discovery of new principles or new scientific laws—has been limited. This article reviews how LLMs are currently used as technical tools to augment the scientific process, and how they might be used in the future as they become more powerful and evolve into robust scientific assistants. By combining data-driven techniques with symbolic systems, such systems can merge into hybrid engines, potentially leading new research directions. We aim to describe the gap between LLMs as technical tools and as "creative engines" capable of fostering high-quality scientific discoveries and posing new questions and hypotheses to human scientists. We first review current applications of LLMs in science, aiming to identify limitations that need to be addressed to move toward creative engines.

Current Applications

LLMs have been widely used in scientific literature analysis, experimental protocol generation, data parsing, and preliminary hypothesis formation. For example, in drug discovery, LLMs can help predict molecular properties; in materials science, they can recommend synthesis pathways. However, most of these applications remain at the assistive level and are not yet capable of independently generating original scientific insights.

Challenges and Limitations

Core challenges facing LLMs include a lack of causal understanding of the physical world, a tendency to "hallucinate," reliance on the quality and biases of training data, and insufficient cross-domain reasoning abilities. In fundamental science, discovering entirely new principles requires abstract reasoning capabilities that go beyond pattern matching—an area where current LLMs are deficient.

Future Directions

To achieve the leap from tools to creative engines, it is necessary to develop hybrid systems that combine LLMs' linguistic capabilities with symbolic reasoning, knowledge graphs, and physical simulations. In addition, establishing clear evaluation metrics is crucial to ensure that LLM outputs align with human scientific goals. Interdisciplinary collaboration and ongoing human-machine alignment will be key.

ConclusionLLMs have the potential to become powerful partners in the scientific method, but to truly drive innovation from hypothesis to discovery, systematic improvements in technology, evaluation, and ethics are still needed.

Tech Logic