Principles of Large Language Models (LLM)

Authors

DOI:

https://doi.org/10.17721/2706-9699.2025.1.06

Keywords:

large language models, transformer, autoregressive generation, tokenization, temperature, probabilistic text modeling

Abstract

This paper explores the operational principles of large language models (LLMs), focusing in particular on the mechanism of next-token generation within the process of autoregressive modeling. It outlines the theoretical foundations of neural language models, the transformer architecture with its self-attention mechanism, and the roles of tokenization and embedding in forming the input representation of text. The study analyzes the main methods for selecting the next token (greedy decoding, top-k sampling, top-p sampling, temperature), their impact on the stochasticity of results, and the trade-off between coherence and creativity. It also examines context length limitations, sources of training data, and challenges related to interpretability and the likelihood of «hallucinations». The article provides a comprehensive overview of the architectural and algorithmic foundations behind text generation in LLMs.

References

Schuurmans D., Dai H., Zanini F. Autoregressive Large Language Models are Computationally Universal. arXiv preprint arXiv:2410.03170 [cs.CL]. 2024.

Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A. N., Kaiser L., Polosukhin I. Attention Is All You Need. arXiv preprint arXiv:1706.03762. 2017.

Grubisic D., Cummins C., Seeker V., Leather H. Priority Sampling of Large Language Models for Compilers. arXiv preprint arXiv:2402.18734. 2024.

Schmidt R. M. Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview. arXiv preprint arXiv:1912.05911. 2019.

Vennerod C. B., Kjorran A., Bugge E. S. Long Short-term Memory RNN. arXiv preprint arXiv:2105.06756. 2021.

Chung J., Gulcehre C., Cho K., Bengio Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv preprint arXiv:1412.3555. 2014. Presented at NIPS 2014 Deep Learning and Representation Learning Workshop. https://doi.org/10.48550/arXiv.1412.3555.

Yenduri G. et al. Generative Pre-trained Transformer: A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions. arXiv preprint arXiv:2305.10435. 2023.

Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805. 2018. https://doi.org/10.48550/arXiv.1810.04805.

Sinha K. et al. Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little. arXiv preprint arXiv:2104.06644. 2021. https://doi.org/10.48550/arXiv.2104.06644.

Batsuren K. et al. Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge. arXiv preprint arXiv:2404.13292. 2024.

Kozma L., Voderholzer J. Theoretical Analysis of Byte-Pair Encoding. arXiv preprint arXiv:2411.08671. 2024. https://doi.org/10.48550/arXiv.2411.08671.

Ma H., Chen J., Wang G., Zhang C. Estimating LLM Uncertainty with Logits. arXiv preprint arXiv:2502.00290. 2025. https://arxiv.org/abs/2502.00290.

PEDAL: Prompts based on Exemplar Diversity Aggregated using LLMs. arXiv preprint arXiv:2408.08869. 2024. https://doi.org/10.48550/arXiv.2408.08869.

Yao Y. et al. Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling. arXiv preprint arXiv:2410.03777. 2024.

Ravfogel S., Goldberg Y., Goldberger J. Conformal Nucleus Sampling. arXiv preprint arXiv:2305.02633. 2023. (Findings of ACL 2023). https://doi.org/10.48550/arXiv.2305.02633.

Renze M., Guven E. The Effect of Sampling Temperature on Problem Solving in Large Language Models. Findings of the Association for Computational Linguistics: EMNLP. 2024. P. 7346–7356. https://doi.org/10.18653/v1/2024.findings-emnlp.432.

Lou C., Jia Z., Zheng Z., Tu K. Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers. arXiv preprint arXiv:2406.16747. 2024.

Beltagy I., Peters M. E., Cohan A. Longformer: The Long-Document Transformer. arXiv preprint arXiv:2004.05150. 2020.

Choromanski K. et al. Rethinking Attention with Performers. arXiv preprint arXiv:2009.14794. 2020. https://doi.org/10.48550/arXiv.2009.14794.

Bulatov A., Kuratov Y., Burtsev M. Recurrent Memory Transformer. arXiv preprint arXiv:2207.06881. 2022. Version 2. https://doi.org/10.48550/arXiv.2207. 06881.

Lewis P. et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv preprint arXiv:2005.11401. 2020.

Published

2025-07-17

How to Cite

Lysyi, P. O. (2025). Principles of Large Language Models (LLM). Journal of Numerical and Applied Mathematics, 1, 63-76. https://doi.org/10.17721/2706-9699.2025.1.06