Principles of Large Language Models (LLM)
DOI:
https://doi.org/10.17721/2706-9699.2025.1.06Keywords:
large language models, transformer, autoregressive generation, tokenization, temperature, probabilistic text modelingAbstract
This paper explores the operational principles of large language models (LLMs), focusing in particular on the mechanism of next-token generation within the process of autoregressive modeling. It outlines the theoretical foundations of neural language models, the transformer architecture with its self-attention mechanism, and the roles of tokenization and embedding in forming the input representation of text. The study analyzes the main methods for selecting the next token (greedy decoding, top-k sampling, top-p sampling, temperature), their impact on the stochasticity of results, and the trade-off between coherence and creativity. It also examines context length limitations, sources of training data, and challenges related to interpretability and the likelihood of «hallucinations». The article provides a comprehensive overview of the architectural and algorithmic foundations behind text generation in LLMs.
References
Schuurmans D., Dai H., Zanini F. Autoregressive Large Language Models are Computationally Universal. arXiv preprint arXiv:2410.03170 [cs.CL]. 2024.
Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A. N., Kaiser L., Polosukhin I. Attention Is All You Need. arXiv preprint arXiv:1706.03762. 2017.
Grubisic D., Cummins C., Seeker V., Leather H. Priority Sampling of Large Language Models for Compilers. arXiv preprint arXiv:2402.18734. 2024.
Schmidt R. M. Recurrent Neural Networks (RNNs): A Gentle Introduction and Overview. arXiv preprint arXiv:1912.05911. 2019.
Vennerod C. B., Kjorran A., Bugge E. S. Long Short-term Memory RNN. arXiv preprint arXiv:2105.06756. 2021.
Chung J., Gulcehre C., Cho K., Bengio Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv preprint arXiv:1412.3555. 2014. Presented at NIPS 2014 Deep Learning and Representation Learning Workshop. https://doi.org/10.48550/arXiv.1412.3555.
Yenduri G. et al. Generative Pre-trained Transformer: A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions. arXiv preprint arXiv:2305.10435. 2023.
Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805. 2018. https://doi.org/10.48550/arXiv.1810.04805.
Sinha K. et al. Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little. arXiv preprint arXiv:2104.06644. 2021. https://doi.org/10.48550/arXiv.2104.06644.
Batsuren K. et al. Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge. arXiv preprint arXiv:2404.13292. 2024.
Kozma L., Voderholzer J. Theoretical Analysis of Byte-Pair Encoding. arXiv preprint arXiv:2411.08671. 2024. https://doi.org/10.48550/arXiv.2411.08671.
Ma H., Chen J., Wang G., Zhang C. Estimating LLM Uncertainty with Logits. arXiv preprint arXiv:2502.00290. 2025. https://arxiv.org/abs/2502.00290.
PEDAL: Prompts based on Exemplar Diversity Aggregated using LLMs. arXiv preprint arXiv:2408.08869. 2024. https://doi.org/10.48550/arXiv.2408.08869.
Yao Y. et al. Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling. arXiv preprint arXiv:2410.03777. 2024.
Ravfogel S., Goldberg Y., Goldberger J. Conformal Nucleus Sampling. arXiv preprint arXiv:2305.02633. 2023. (Findings of ACL 2023). https://doi.org/10.48550/arXiv.2305.02633.
Renze M., Guven E. The Effect of Sampling Temperature on Problem Solving in Large Language Models. Findings of the Association for Computational Linguistics: EMNLP. 2024. P. 7346–7356. https://doi.org/10.18653/v1/2024.findings-emnlp.432.
Lou C., Jia Z., Zheng Z., Tu K. Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers. arXiv preprint arXiv:2406.16747. 2024.
Beltagy I., Peters M. E., Cohan A. Longformer: The Long-Document Transformer. arXiv preprint arXiv:2004.05150. 2020.
Choromanski K. et al. Rethinking Attention with Performers. arXiv preprint arXiv:2009.14794. 2020. https://doi.org/10.48550/arXiv.2009.14794.
Bulatov A., Kuratov Y., Burtsev M. Recurrent Memory Transformer. arXiv preprint arXiv:2207.06881. 2022. Version 2. https://doi.org/10.48550/arXiv.2207. 06881.
Lewis P. et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv preprint arXiv:2005.11401. 2020.