Spyra-20B: A Proof-of-Concept for Explicit Architectural Design   Reasoning in Domain-Specialised LLMs

Authors

DOI:

https://doi.org/10.38027/smart.v3n1-2

Keywords:

AI, Architecture, Large Language Models, Tree-of-Thought, Chain-of-Thought, Design Reasoning

Abstract

Domain-specifically fine-tuned language models (BloombergGPT, Med-PaLM) have proven a viable route to adapting vocabulary and specialist knowledge to particular fields of application. For architectural practice, however, the central challenge is not terminology but the deliberative structure of the design decision: the weighing of competing objectives across building law, quality of use, cost and sustainability. This paper presents Spyra-20B, a proof-of-concept produced by QLoRA fine-tuning of gpt-oss-20B on 1,693 architecture- and planning-specific dialogue examples. Two features characterise the approach: a three-level deliberation depth controllable by metadata (reasoning_effort), and a strict separation of analysis and answer channels that renders the deliberative process inspectable. An evaluation on 50 law- and planning-related items across four conditions (base model vs. Spyra-20B, closed-book vs. open-book with RAG; 4,000 requests) yields a negative headline result: the fine-tuned model attains roughly six percentage points lower answer accuracy (Spyra 61.1% / 70.4%, base 67.3% / 76.3%; p < 0.001), while retrieval improves both models by about nine percentage points. Fine-tuning changes the form of the output: more schematic, twice as fast at the median, but with a sixfold increase in unrecoverable format failures. We derive four governance requirements: reasoning transparency, external factual grounding, local executability and multi-sample verification, and conclude that such systems should be assessed along architectural properties rather than scores on individual benchmarks.

References

Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., … Zoph, B. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774

Bashir, A. H., Khalid, M. R., Cvejoski, K., Birr, J., Berghaus, J., Berger, A., Halscheidt, S., Temath, C., Sifa, R., & Berghaus, D. (2026). Domain-adaptation through synthetic data: Fine-tuning large language models for German law. arXiv. https://doi.org/10.48550/arXiv.2601.14160

Bommarito, M., II, & Katz, D. M. (2022). GPT takes the bar exam. arXiv. https://doi.org/10.48550/arXiv.2212.14402

Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., … Fiedel, N. (2023). PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240), 1–113. https://doi.org/10.5555/3648699.3648939

Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36, 10088–10115. https://doi.org/10.52202/075280-0441

Fernandes, R., Biedenkapp, A., Hutter, F., & Awad, N. (2025). A Llama walks into the ‘Bar’: Efficient supervised fine-tuning for legal reasoning in the multi-state bar exam. arXiv. https://doi.org/10.48550/arXiv.2504.04945

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv. https://doi.org/10.48550/arXiv.2312.10997

Hirsekorn, Y., Ansre, N., & Grunwald, G. (2025). AI-driven knowledge transfer in architectural education. Smart Design Policies, 2(1), 122–139. https://doi.org/10.38027/smart.v2n1-8

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2106.09685

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), Article 42. https://doi.org/10.1145/3703155

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730

Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2022). Large language models are zero-shot reasoners. Advances in Neural Information Processing Systems, 35, 22199–22213. https://doi.org/10.52202/068431-1613

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. https://doi.org/10.5555/3495724.3496517

Li, P., Li, B., & Li, Z. (2023). Sketch-to-architecture: Generative AI-aided architectural design. In Pacific Graphics short papers and posters (pp. 99–102). Eurographics Association. https://doi.org/10.2312/PG.20231276

Li, Z., Xia, L., Tang, J., Xu, Y., Shi, L., Xia, L., Yin, D., & Huang, C. (2024). UrbanGPT: Spatio-temporal large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5351–5362). Association for Computing Machinery. https://doi.org/10.1145/3637528.3671578

Ling, C., Zhao, X., Lu, J., Deng, C., Zheng, C., Wang, J., Chowdhury, T., Li, Y., Cui, H., Zhang, X., Zhao, T., Panalkar, A., Mehta, D., Pasquali, S., Cheng, W., Wang, H., Liu, Y., Chen, Z., Chen, H., … Zhao, L. (2025). Domain specialization as the key to make large language models disruptive: A comprehensive survey. ACM Computing Surveys, 58(3), Article 79, 1–39. https://doi.org/10.1145/3764579

Madireddy, S., Gao, L., Din, Z., Kim, K., Senouci, A., Han, Z., & Zhang, Y. (2025). Large language model-driven code compliance checking in building information modeling. Electronics, 14(11), 2146. https://doi.org/10.3390/electronics14112146

Meyerson, E., Paolo, G., Dailey, R., Shahrzad, H., Francon, O., Hayes, C. F., Qiu, X., Hodjat, B., & Miikkulainen, R. (2025). Solving a million-step LLM task with zero errors. arXiv. https://doi.org/10.48550/arXiv.2511.09030

Nannini, L., Balayn, A., & Smith, A. L. (2023). Explainability in AI policies: A critical review of communications, reports, regulations, and standards in the EU, US, and UK. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (pp. 1198–1212). Association for Computing Machinery. https://doi.org/10.1145/3593013.3594074

OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774

OpenAI. (2025). gpt-oss-120b & gpt-oss-20b model card. arXiv. https://doi.org/10.48550/arXiv.2508.10925

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. https://doi.org/10.52202/068431-2011

Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Agüera y Arcas, B., … Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172–180. https://doi.org/10.1038/s41586-023-06291-2

Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv. https://doi.org/10.48550/arXiv.2307.09288

Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-consistency improves chain-of-thought reasoning in language models. International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2203.11171

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. https://doi.org/10.52202/068431-1800

Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., … Gabriel, I. (2021). Ethical and social risks of harm from language models. arXiv. https://doi.org/10.48550/arXiv.2112.04359

Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A large language model for finance. arXiv. https://doi.org/10.48550/arXiv.2303.17564

Yang, F., & Zhang, J. (2024). Prompt-based automation of building code information transformation for compliance checking. Automation in Construction, 168, Article 105817. https://doi.org/10.1016/j.autcon.2024.105817

Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 11809–11822. https://doi.org/10.52202/075280-0517

Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., Wang, L., Luu, A. T., Bi, W., Shi, F., & Shi, S. (2025). Siren’s song in the AI ocean: A survey on hallucination in large language models. Computational Linguistics, 51(4), 1373–1418. https://doi.org/10.1162/coli.a.16

Downloads

Published

2026-08-15

How to Cite

Spyra-20B: A Proof-of-Concept for Explicit Architectural Design   Reasoning in Domain-Specialised LLMs. (2026). Smart Design Policies, 3(1), 14–28. https://doi.org/10.38027/smart.v3n1-2

Share

Most read articles by the same author(s)

Similar Articles

11-20 of 27

You may also start an advanced similarity search for this article.