Large Language Models: Architectures and Applications
DOI:
https://doi.org/10.5281/zenodo.19614627Keywords:
large language models; GPT; mixture of experts; retrieval augmented generation; RLHF alignment; instruction tuning; emergent abilities; scaling lawsAbstract
Large language models have rapidly transitioned from research curiosities to foundational infrastructure underpinning applications across virtually every domain that involves natural language. This study presents a systematic empirical comparison of five LLM architectural paradigms -- dense decoder-only (GPT-style), sparse mixture-of-experts (MoE), encoder-decoder (T5-style), retrieval-augmented (RETRO-style), and instruction-tuned with RLHF alignment -- across six application categories: text generation quality (summarisation, translation), reasoning (arithmetic, logical, commonsense), code generation (HumanEval, MBPP), knowledge-intensive QA (TriviaQA, Natural Questions), instruction following (AlpacaEval), and safety/toxicity (ToxiGen, RealToxicityPrompts). Models were evaluated at three parameter scales (1.3B, 6.7B, and 13B) under matched training compute (approximately 300B tokens each) to isolate architectural effects from scale effects. A total of 2,880 evaluations were conducted. Dense decoder-only models achieved the strongest generation quality (ROUGE-L = 42.8 +- 0.6 on CNN/DailyMail at 13B) and code generation (HumanEval pass@1 = 28.4 +- 1.2% at 13B). MoE models achieved equivalent quality at 2.4x lower inference FLOPs by activating only 25% of parameters per token. Retrieval-augmented models reduced hallucination by 38.4% on knowledge QA but added 1.8x latency from retrieval. RLHF alignment improved instruction following (AlpacaEval win rate 74.2% vs. 52.8% for base model) and reduced toxicity by 64.2% but introduced a 2.8% regression on reasoning benchmarks (the alignment tax). Scaling from 1.3B to 13B improved reasoning most steeply (+18.4 pp on arithmetic) while generation quality improved modestly (+4.2 pp ROUGE-L), suggesting that reasoning benefits disproportionately from scale. A practical LLM selection framework mapping application requirements to recommended architecture and scale is proposed.Downloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
Large Language Models: Architectures and Applications. (2026). Bio-QI Journal, 2(1), 1-9. https://doi.org/10.5281/zenodo.19614627

