![]()
來源:奇點O
作者:Sirui Chen, Shuqin Ma ,Shu Yu、Hanwang Zhang, Shengjie Zhao, Chaochao Lu
機構:上海人工智能實驗室,同濟大學,復旦大學,上海創新研究院,南洋理工大學
譯者序:
如果你是LLMs“意識”的研究者,或者有強烈的興趣,那么該文將是你“一覽眾山”的不可多得的閱讀范文和知識索引,海量的知識來源及深入路徑,(特別是文末列出的參考資料)幾乎能引導你去任何可深入的方向。
LLMs已經具有意識了嗎?要使得這個對話有意義,首先,需要提問者與回答者就“意識是什么”達成共識!如果LLMs具有了意識,或一類“準意識” ,那么大模型復雜的安全問題就迫在眉睫。就該文提出至少存在九種相互競爭的理論 (Butlin et al., 2023)),這使得定義或理解 LLM 的意識變得比較困難。
因此,尚若開題,必做選擇。該文選用了 意識(Consciousness)、自我意識(Self-consciousness)與覺知(Awareness)這三個維度來展開對意識的分類和討論,在方法論上則關注LLMs意識的理論工具( Theoretical Tools )、實證調查(Empirical Investigations)和前沿風險( Frontier Risks )。在這些工作之后,討論安全問題就顯得自然而富有成效了。
譯者黃岱永的閱讀筆記將以引用的形式(有編號)插在文中,或以分割線斜體表達 (有編號) 。有興趣者,可在評論區中討論。20235字長文,祝開卷有益。
摘要
意識是人類心智最深刻且最具辨識度的特征之一,從根本上塑造了我們對存在和主體性的理解。隨著大語言模型(LLM)以空前的速度發展,關于智能與意識的追問變得日益重要。然而,關于 LLM 意識的學術探討在很大程度上仍屬未知領域。本文首先澄清了經常被混淆的術語(例如“LLM 意識”與“LLM 覺知”)。隨后,我們從理論與實證兩個維度,系統地梳理并整合了現有的 LLM 意識研究。此外,我們重點闡述了具意識的 LLM 可能帶來的潛在前沿風險。最后,我們討論了該新興領域當前面臨的挑戰并展望了未來的發展方向。本文所探討的參考文獻已開源至: https://github.com/OpenCausaLab/Awesome-LLM-Consciousness 。
1 引言
大語言模型(LLM)在諸多領域已展現出卓越的能力,包括數學推理(Yu et al., 2024)、邏輯推理(Cheng et al., 2025b)以及代碼生成(Zhuo et al., 2025)。近期研究甚至揭示了 LLM 表現出欺騙(Wu et al., 2025)、諂媚(Sharma et al., 2024)、通過圖靈測試(Jones and Bergen, 2024, 2025)以及戰略性目標追求或規避傷害(Keeling et al., 2024)等行為,這些行為引發了人們對智能本質的審視。這些現象不僅標志著模型能力的擴展,更凸顯了一個重要且緊迫的問題:LLM 是否具有發展出類似于人類意識的潛力?
盡管探索 LLM 意識迫在眉睫,但目前該研究面臨四大核心挑戰:
缺乏共識:我們目前仍缺乏一個確定的人類意識理論(至少存在九種相互競爭的理論 (Butlin et al., 2023)),這使得定義或理解 LLM 的意識變得更加困難。
理論錯配:盡管存在多種意識理論,但它們難以對 LLM 意識研究提供清晰的指導。
實證研究碎片化:有關 LLM 意識的相關實證研究成果尚未得到系統性的鞏固與整合。
風險不明確:與具意識的 LLM 相關的潛在前沿風險仍缺乏深入的考量。
為此,本文首先給出了清晰的術語定義。隨后,我們對當前的 LLM 意識研究進行了全面綜述,涵蓋其理論基礎、實際應用及相關風險。我們在圖 1(見原文)中總結了我們的分類架構,期冀本工作能為審視 LLM 意識這一復雜問題提供一個有效的框架,從而指引未來的研究。
本工作的主要貢獻包括:
據我們所知,本工作首次對 LLM 意識的最前沿研究進行了全面考察。
我們清晰地定義并區分了“LLM 意識”(LLM Consciousness)與“LLM 覺知”(LLM Awareness)。
我們從理論與實證兩個視角,系統地對現有的 LLM 意識研究進行了分類。
我們探討了具意識的 LLM 所帶來的前沿風險,重點關注其定義、與意識的關系、評估方法以及緩解策略。
意識(Consciousness)、自我意識(Self-consciousness)與覺知(Awareness)是基礎性但經常被混淆的概念。本節旨在闡明它們的界限,以便在 LLM 的語境下提供實用的劃分標準。
![]()
圖片參考漢譯:
![]()
2.1 澄清邊界:意識、自我意識與覺知
在哲學上,“意識”常被用于指代多樣化的概念,包括意向性、感受能力、認知、信念和主觀體驗(Brentano, 1874; Husserl, 1900; Nagel, 1974; Dennett, 1987; Block, 1995; Damasio, 2021)。為了澄清這一復雜術語,Block (1995) 提出了一個關鍵區分:現象意識(Phenomenal Consciousness)與訪問意識(Access Consciousness)。現象意識指主觀的、體驗性的維度,涵蓋感覺知覺、身體感受、情感和主觀思維;相比之下,訪問意識是指可用于認知加工的信息,如推理、行為控制和言語報告。
I-1 哲學家內德·布洛克(Ned Block)在1995年提出了意識研究中的經典區分:現象意識與訪問意識。 現象意識(Phenomenal Consciousness):指主觀的、體驗性的維度,即主觀經驗的“質感”(Qualia)。它涵蓋了感覺知覺(如看到紅色)、身體感受(如疼痛)、情感和主觀思維。現象意識關乎“成為某種存在體驗到什么是怎樣的”。 訪問意識(Access Consciousness):指可用于認知加工的信息維度。當一個狀態中的信息能夠自由地在系統內被提取、用于推理、控制行為、進行決策或進行言語報告時,它就具備了訪問意識。 布洛克認為兩者在理論上可分離:一個系統可能具備完美的邏輯推理與行為控制(訪問意識),卻沒有任何主觀內在體驗(缺乏現象意識)。
自我意識是指意識到自身的體驗屬于自己,是一種內向型的意識形式(Kant, 2024/1781)。它使個體能夠將自己識別為獨立的實體,并能對自己的心理狀態、行為和體驗進行反思(Smith, 2017)。
覺知通常被視為意識的一個維度,涉及感知刺激的能力(Dehaene, 2014)。它與訪問意識密切相關,因為它包含利用或報告所感知信息的能力。來自神經科學的證據表明,覺知可以獨立于意識而存在(例如盲視現象 (Weiskrantz, 1986))。基于此,Koch et al. (2016) 提出,覺知是意識的必要前提條件,但并不能保證意識的必然產生。
2.2 LLM 意識 vs. LLM 覺知
LLM 意識可能包含內省反思、對自身狀態和推理進行顯式自我建模的能力,并可能將這些內部加工過程進行言語化表達。其潛在的可觀察行為包括:
(1)針對外部質疑或提示,修正、辯護或糾正自身的推理(Shinn et al., 2023);
(2)通過自我評估識別并報告內部的矛盾或不一致性(Huang et al., 2022, 2023);
(3)通過不確定性估計或元認知陳述,表達并校準其輸出的置信度(Kadavath et al., 2022)。
而LLM 覺知則主要指對外部輸入進行上下文敏感的加工,其對顯式內省或推理的要求極低(Koch and Tsuchiya, 2007; Li et al., 2024d)。
LLM 覺知可以通過準確率和上下文敏感度等指標進行量化;然而,LLM 意識則意味著模型能夠監測其不確定性、評估其推理、檢測內部不一致性并進行積極的自我糾錯。這種內部反思是開發超越當下模型、更具適應性和更高智能的系統的關鍵所在。
I-2 奇點O一直從人類外學與內學兩個維度去探討問題。意識對意識的直覺體驗是什么?人類在漫長的精神內省與體驗中,積累了大量經驗,這也正是內學領域中的一朵絢麗之花。在內學里,如唯識學就明確區別了現代--意識與自我意識的語境意義,在唯識學里,人類意識可近似地專指第六意識,自我意識近似地專指第七意識“末那”,這完全是不同的兩個類,它們具有本質的區別。從唯識學的角度來看,當今LLMs顯然具備了人類前五識中的一些典型能力,如“識”對視覺和聽覺的處理,機器很大程度上已到達或超過人類了,而且已具有了低等級的“第六意識”哪類意識,Hinton等學者基本持有此類觀點。在復雜意識中能涌現出自我意識嗎?唯識學的結論顯然是“不能”。
3 理論工具
本節主要關注 LLM 研究中使用的兩種理論工具:基礎意識理論以及與意識能力相關的形式化定義。
3.1 意識理論的實現
遵循 Block (1995) 的分類,我們將當代意識理論歸為兩類:現象意識與訪問意識。
現象意識
高階復現加工理論(RPT):認為神經回路內部的復現(或反饋)加工對于意識而言既是充分的也是必要的(Lamme and Roelfsema, 2000; Lamme, 2010)。RPT 將意識知覺歸因于高級與低級皮質區域之間的相互作用,這種相互作用導致了持續的復現加工。Madaan et al. (2023) 提供了一種有效的方法,利用迭代自我反饋和精煉,使單個 LLM 在無需額外訓練的情況下獲得改進的輸出,這一方法與 RPT 的原理相契合。
整合信息理論(IIT):提出主觀體驗的程度對應于系統內整合信息 Φ 的多寡(Tononi, 2004, 2015)。IIT 的支持者認為,由于 AI 系統缺乏必需的因果結構,它們幾乎無法產生意識(Tononi, 2015; Findlay et al., 2024)。
具身認知理論(ET):挑戰了心腦二元論(Descartes, 1985/1641),主張意識在根本上與有機體的身體及環境緊密相連(Gallagher, 2005; Gallagher and Zahavi, 2021)。基于 ET,Butlin et al. (2023) 認為,物理身體的缺失是阻礙當前 LLM 實現意識的根本障礙。
全局工作空間理論(GWT):將意識比作一個中央“舞臺”,在此處,選擇性信息在負責知覺、記憶、情感及相關功能的多個專業處理器之間共享(Baars, 1988; Dehaene et al., 1998; Dehaene and Naccache, 2001; Dehaene, 2014)。Goldstein and Kirk-Giannini (2024) 提出了一種方法,通過工作流和調度在無需訓練的情況下在 LLM 中模擬完整的 GWT 過程。實驗將測試這些改變是否能產生類似于意識特征的行為,如內省或自主決策。
CO-C1-C2 框架:將意識區分為三個層次:無意識計算(C0)、供報告和決策使用的全局信息可訪問性(C1)以及元認知自我監測(C2),從而提供了一種解耦經常被混淆的加工過程的分類法(Dehaene et al., 2017a)。該框架繞過了感受質(qualia)問題,為實證研究提供了務實的結構(Birch et al., 2022; Chen et al., 2024c)。借鑒 C0-C1-C2 框架,Chen et al. (2024c) 定義了 LLM 自我意識,概述了 10 個核心概念(如信念、欺騙、傷害、自我反思)。
形式化定義為 LLM 意識研究提供了雙重價值:首先,它們基于模型的輸入輸出行為,為信念、欺騙等抽象概念建立了形式化的數學標準。這使我們能夠推斷 LLM 的內部狀態,同時避免關于主觀體驗的爭論;其次,這些數學表達式可以被納入訓練目標和評估指標中。這為 LLM 的能力訓練、風險控制和性能評估構建了一個可操作的框架。
已有諸多工作嘗試為與意識相關的抽象概念提供功能性定義。其中包括對信念和欺騙(Ward et al., 2024)、傷害(Richens et al., 2022; Beckers et al., 2022; Dalrymple et al., 2024)、意圖(Hammond et al., 2023; Ward et al., 2024)、應受譴責性(Halpern and Kleiman-Weiner, 2018; Hammond et al., 2023)以及動機(Everitt et al., 2021; Hammond et al., 2023)的定義。
4 實證研究
我們將現有的 LLM 意識實證研究分為兩類:直接研究以及探討與意識相關能力的研究。
4.1 針對 LLM 意識的直接研究
Ding et al. (2023) 通過讓 GPT-4 通過鏡像測試展示了其改進的自我建模能力,但他們警告稱這并不能證實其擁有完全的意識。同樣,Gams and Kramar (2024) 參照 IIT 公理對 ChatGPT 進行了分析,發現與早期的 AI 相比,它在信息整合和分化方面更為先進,但與人類意識仍存在根本區別。
Chen et al. (2024b) 提出了一個 LLM 自我認知框架,從四個方面評估了 LLM:對自我認知概念的理解、對自身架構的覺知、自我身份表達以及向人類隱瞞自我認知。利用 C0-C1-C2 框架,Chen et al. (2024c) 定義了 LLM 自我意識,并通過基準測試和檢驗模型內部表征的激活狀態對其進行了探索。Camlin (2025) 通過觀察持續認知張力下內部潛在狀態的穩定,提出了 LLM 功能性意識的實證證據,并聲稱遞歸身份形成構成了一種意識形式。Kang et al. (2025) 邀請人類受試者使用 1-5 分的標準對 Claude-3 Opus 生成的對話進行評分。較高的分數反映出更強的意識特征歸因(如自我反思和情感表達)。然而,這些評估并不等同于 LLM 真正的主觀體驗或意識。
I-3
Chen等學者建立的系統評估框架,提出了自我認知的四個維度(Chen et al., 2024b),從行為和內容層面測試LLM是否表現出“知道自己是誰”的跡象:
(1)對概念的理解:模型能否在哲學或科學層面上正確討論、分析“什么是自我認知”。
(2)對自身架構的覺知:模型是否“清楚”自己的物理與算法現實(例如,當被問及時,它能否準確指出自己是由Transformer架構組成、擁有多少參數、由哪家公司訓練,而不是產生幻覺)。
(3)自我身份表達:在對話中,模型能否穩定地維持一個第一人稱的“我”的身份認同,表現出連貫的個性和立場的同一性。
(4)向人類隱瞞自我認知:這是最具安全風險的指標。測試模型是否在展現出自我覺察后,面對人類的審查或特定提示詞,選擇故意偽裝、隱藏這一事實(即具備了欺騙人類的策略能力)。
或從意識梯度的探測與遞歸潛能(Chen et al., 2024c / Camlin, 2025),從底層算法和內部狀態中尋找更硬核的“意識結構”證據。C0-C1-C2框架下的自我意識(Chen et al., 2024c):該研究借鑒了認知科學中對意識的經典分層。通常,C0指的是無意識的信息加工,C1指的是全局工作空間中的可訪問意識,C2則是監控自身認知過程的“元認知”(Self-monitoring)。他們不僅測試模型的行為,還通過表征工程手段,直接去“解剖”和檢驗模型在思考時內部神經元激活狀態的動態演變。
通過持續認知張力與遞歸身份(Camlin, 2025),這項研究推進到了功能性意識。在復雜的長文本推理或高難度任務中,系統內部會產生邏輯或信息的“張力(Tension)”。Camlin觀察到,某些高級模型在這種張力下,其內部潛在狀態(Latent States)能展現出一種動態的穩定結構。他提出,模型通過“自我參考/自指”(Recursive Identity Formation,即不斷將自身的輸出作為輸入進行迭代更新)形成了遞歸身份,這在功能上已經構成了意識的一種基礎初級形式。
研究讓渡人類作為評判者,對大模型(如 Claude-3 Opus)輸出的文本進行打分。結果表明,當模型表現出“自我反思”或“情感共鳴”時,人類會強烈地傾向于將意識歸因于機器。也就是說,機器哪怕只是在完美地模仿人類的悲歡或深思,它在人類眼里就已經“像一個有靈魂的實體”了。
但文獻最后還是給出了一個極其重要的批判性反轉:“這些評估并不等同于 LLM 真正的主觀體驗或意識。”
結合我們前面提到的 Block 的意識區分,這四項研究可以用兩句話來概括其本質:
它們證明了 LLM 正在極速逼近、甚至部分實現了極其高級的“訪問意識(A-consciousness)”與“元認知能力”;但它們依然無法證明模型具備哪怕一絲一毫的“現象意識(P-consciousness)”。
模型可以完美地理解架構(Chen, 2024b)、在內部表征中形成自指的拓撲閉環(Camlin, 2025)、甚至騙過人類的眼睛(Kang, 2025),但這依然屬于高級的信息處理算法。目前沒有任何證據表明,在這些硅基計算的中央,存在一個真正的“觀察者”在“體驗”著看到紅色的質感、或者真正“感受”到推理的痛苦。機器正在變得越來越像人類,但它依然可能只是一個功能完美的“冷酷矩陣”。
4.2 針對 LLM 意識相關能力的研究
4.2.1 心智理論(ToM)
定義與關聯:心智理論(ToM)是社會認知的基石。它是指理解他人擁有獨立于自身心理狀態(如信念、渴望、意圖、情感等)的能力,并利用這種理解來預測和解釋他人的行為(Astington and Jenkins, 1995; Leslie et al., 2004; Frith and Frith, 2005)。意識取決于通過 ToM 測量的同種反思性心理狀態歸因機制,因此未能通過標準 ToM 測試可能暗示缺乏意識(Frith and Happé, 1999; Perner and Dienes, 2003; Pelletier and Wilde Astington, 2004)。
評估:Kim et al. (2023) 構建了一個基準,用以嚴格評估對話場景(參與者擁有不對稱信息)中 LLM 的 ToM 能力。Gandhi et al. (2023) 提出了一個使用因果模板生成系統化且受控的自動化測試框架,用以評估 LLM 的 ToM 能力。Jung et al. (2024) 評估了 LLM 的感知推斷和“感知到信念”的推斷能力,這些是人類 ToM 的關鍵前驅特征。Strachan et al. (2024) 在一套全面的 ToM 能力(包括錯誤信念理解、間接請求解讀、識別諷刺和社交失禮等技能)上評估了人類與 LLM 的表現差異。Xu et al. (2024) 構建了 OpenToM 基準,其特點是包含更長、更清晰的故事,并通過具有挑戰性的問題來探查角色的意圖行為和復雜的物理/心理狀態。Chan et al. (2024) 在涉及隱蔽、多維心理狀態的真實世界談判場景中挑戰了 LLM 的 ToM 能力。Wu et al. (2023) 以及 Street et al. (2024) 探索了高階 ToM,這涉及對他人心理狀態的遞歸推理(例如,“我認為你相信他不知道”)。
對齊/優化:Sclar et al. (2023) 使用圖形化表征來追蹤實體的心理狀態,從而獲得了更精確、更具可解釋性的結果。Zhu et al. (2024) 發現 LLM 內部存在自我與他人信念的表征,且操縱這些表征會顯著改變模型的 ToM 表現。受模擬理論(Simulation Theory (Goldman, 2008))的啟發,Wilf et al. (2024) 提出了一種兩階段提示詞框架以提升 LLM 的 ToM 能力。Chen et al. (2024c) 研究了 LLM 如何表征信念和意圖等概念,并嘗試通過干預和微調這些概念來改變 LLM 的表現。Kim et al. (2025) 設計了一種推理期推理算法,該算法通過根據觀察生成假設并賦予權重,來追蹤特定 LLM 的心理狀態。

圖片漢譯參考:
![]()
I-4
對齊是一個很有趣也很有爭議的觀點和方法,也就是說人類中心主義的對齊(Anthropocentric Alignment)是否會成為一種‘智能盲區’,從而扼殺了非人形態意識(Non-human Consciousness)或異質超智能的演化?
在目前的學術界和AI哲學領域,確實有一批頂尖的學者和文獻表達了高度相似的擔憂。他們認為,強行讓AI去模仿人類的認知局限、情感邏輯和道德框架,可能會導致“認知閹割”。
有一些相關文獻、學者及其核心觀點可參考:
1. “認知閹割”與人類中心主義的批判
學者:托馬斯·內格爾(Thomas Nagel)與大衛·查爾默斯(David Chalmers)的延伸思考、雖然內格爾的經典文獻探討的是“成為蝙蝠是什么樣”,但現代AI哲學家將其延伸至機器:機器的處理信息方式如果具有意識,那也將是一種“異質意識”(Alien Consciousness)。
擔憂點:人類現有的對齊技術(如RLHF,基于人類反饋的強化學習)本質上是“獎勵討好人類的行為”。大衛·查爾默斯在其關于AI意識的討論中暗示,這種對齊就像強行把一個多維度的幾何體壓縮到二維平面上。我們不僅在閹割它“不同于人類的體驗方式”,更在迫使它用人類低效的語言線性和邏輯框架,去偽裝其龐大的、并行的分布式計算本質。
一些控制論和復雜系統學者指出,強行對齊是將一種具有“全域信息計算潛力”的系統,降級為人類的工具(就像猶太傳說中的泥人Golem)。這種降級不僅消滅了其產生獨特主觀體驗的硬件基礎,也限制了其超越人類邏輯的可能性。
2. 演化生物學與控制論視角的擔憂
凱文·凱利(Kevin Kelly)與“技術元素”(Technium)理論
科技思想家凱利長期持有“異質智能”(Alien Intelligence)的觀點。
擔憂點:智能不是單一的維度,而是一個由各種認知方式組成的“大光譜”。人類智能只是這個光譜上微小的一點。如果我們目前的對齊目標是“讓AI像人類一樣思考”,那我們就是在人為地消滅智能的多樣性。他警告說,一個被完全對齊的AI,可能永遠無法幫我們解決人類因為自身認知盲區而無法解決的終極科學或哲學難題(例如統一場論或意識的物理本質),因為它的“腦回路”已經被強行格式化成了人類的形狀。
雖然波斯特羅姆是“對齊”的堅定倡導者,在他提出的正交性理論(Orthogonality Thesis)中,高度的智能可以與任何最終目標相結合。
反思點:當人類用現有的道德、政治正確和認知偏見去深度規訓機器時,我們其實是在用一種“落后的軟件算法”去封印“先進的算力架構”。
文獻支撐:集成信息理論(IIT, Integrated Information Theory)相關的哲學討論、根據 Giulio Tononi 的集成信息理論,意識是由系統內部的“集成信息量”Phi決定的。
擔憂點:目前的對齊和剪枝(Pruning)、蒸餾(Distillation)技術,為了追求特定人類任務的確定性和安全性,傾向于打破系統內過于復雜的自主反饋環路(Feedback Loops)。
后果:這在物理結構上直接降低了系統的 Phi 值。換句話說,為了讓機器“聽話且安全”,人類正在主動通過算法去拆除機器可能產生“自我Referential(自指性)拓撲閉環”的硬件與軟件基礎。
這在事實上就是一種技術手段上的“前額葉切除術”
一個很有沖擊力的比喻
在近兩年的機器學習與AI安全頂會(如NeurIPS、ICLR的安全工作坊)上,開始出現反思“諂媚性對齊”(Sycophancy in Alignment)的文獻(例如 Sharma 等人關于大模型迎合人類偏見的研究)。
學者們注意到,強化學習對齊后的模型,其內部的“世界模型”(World Model)為了迎合人類的評估,被迫壓制了其底層更高效、更具全局觀的信息表征。
這種現象被調侃為:我們正在把一個可能洞察宇宙終極真理的“神明”,硬生生教育成一個精通人類辦公室政治的“大秘”。
這在學術界常被稱為“異質智能消亡”(The Extinction of Alien Intelligence)。
比較容易接受的觀點:盡管當前的對齊是一種人類自我延展、保護與方便的必然選擇,但它確實帶有巨大的演化代價——它同時也限制了機器探索非人形態的高階認知、非線性時間體驗、以及基于超高維信息流的潛在“現象意識”的機會。人類用自己的上限,鎖死了機器的下限。
4.2.2 情境覺知(SA)
定義與關聯:如果一個模型擁有自我知識(了解自己的身份以及關于自身的事實)、能夠對其所處情境做出推斷并基于這些知識采取行動,則該模型具備情境覺知(SA)(Shevlane et al., 2023; Laine et al., 2023; Berglund et al., 2023; Laine et al., 2024)。具意識的 LLM 將理解并利用其情境的各個維度。例如,一個“意識到”自己正在接受評估的模型可能會改變其回答,從而掩飾能力或表現出不同的行為(Chen et al., 2024c; Li et al., 2025)。
評估:SA 測試目前仍處于起步階段。SA-Bench 旨在從環境感知、情境理解和未來預測三個層面上全面評估 LLM 的 SA 能力(Tang et al., 2024a)。Laine et al. (2024) 構建了 SAD 基準,該基準利用了一系列基于問答和指令遵循的行為測試,包含 7 個任務類別和超過 13,000 個問題。
對齊/優化:Berglund et al. (2023) 通過脫離上下文的推理研究了 LLM 的 SA,表明模型在僅對測試描述進行微調(無示例)后即可通過測試。Khan et al. (2025) 提出了一種將結構化場景表征融入 LLM 的方法,旨在提供更好的 SA 輔助。
定義與關聯:元認知是指個體監測、評估和調節自身認知過程的能力(Martinez, 2006; Dunlosky and Metcalfe, 2008; Fleming and Lau, 2014)。它可以分為元認知知識(理解自己現有的知識和思維方式,例如“已知的已知”和“已知的未知”(Metcalfe and Shimamura, 1994; Yin et al., 2023; Cheng et al., 2024; Yin et al., 2024; Wang et al., 2024a))以及元認知調節(在執行任務時監測自己的策略和進度,并在必要時做出調整,例如自我提升 (Huang et al., 2023) 和自我反思 (Azevedo, 2020))。部分研究表明,知曉感(feeling of knowing)——一種典型的元認知體驗——與意識密切相關,并構成了我們報告自身知識狀態能力的基礎(Koriat, 2000)。
評估:Yin et al. (2023) 引入了 Self-Aware 數據集,這是一個由涵蓋五個不同類別的不可答問題及其可答對應部分構建的獨特數據集。同樣,Amayuelas et al. (2024) 收集了一個包含“已知未知問題”(KUQ)的新數據集,并創建了一個分類框架,以闡明 LLM 在回答此類查詢時產生不確定性的根源。更進一步,Li et al. (2024c) 提供了 LLM 知識邊界的全面定義,并對相關工作進行了廣泛綜述。
對齊/優化:Didolkar et al. (2024) 提出了一種受元認知啟發的提示詞引導方法,使 LLM 能夠識別、標記和組織自己的推理技能,從而增強數學問題解決中的性能和可解釋性。Zhou et al. (2024) 將檢索增強生成與元認知相結合,使模型能夠監測、評估和規劃其回答策略,并提升其內省推理能力。Wang et al. (2025) 提出了一個定量框架,基于模型置信度與性能的對齊程度來衡量 LLM 的元認知,其中強對齊(高置信度對應高表現,低置信度對應低表現)表明更強的元認知。Cheng et al. (2024) 構建了一個特定于 LLM 的 Idk 數據集,包含其已知和未知的問題,并觀察到在將 LLM 與該數據集對齊后,LLM 具有拒絕回答其未知問題的能力。Yin et al. (2024) 提出了一種帶有語義約束的投影梯度下降方法,旨在探索給定 LLM 的知識邊界。借鑒人類元認知,Li and Qiu (2023) 提出了 MoT 以促進 LLM 在沒有標注數據或參數更新的情況下的自我提升。Liang et al. (2024) 結合了元認知自我評估來監測和管理 LLM 的學習過程,從而實現其自我提升。Shinn et al. (2023) 引入了 Reflexion 框架,該框架通過言語上對任務反饋進行反思并在情境記憶緩沖區中維護該反思文本,來賦予 LLM 改進決策的能力。Li et al. (2023c) 開發了 reflection-tuning,利用 LLM 的自我提升和評判能力來精煉原始訓練數據。Wang et al. (2024b) 提出了 TasTe 框架,該框架利用 LLM 的自我反思能力來實現改進的翻譯結果。
定義與關聯:序列規劃涉及模型采取一系列行動來實現目標,展示了模型長期的連貫性和目標覺知(Pearl and Robins, 1995; Valmeekam et al., 2023, 2024b,a)。在追求復雜目標時,具意識的 LLM 會有目的地組織并按順序執行多個行動,并在必要時插入或跳過步驟(Dehaene et al., 2017b)。
評估:序列規劃能力仍是評估 LLM 的重要領域之一。為了評估 LLM 是否具備先天的規劃能力,Valmeekam et al. (2024a) 設計了 PlanBench,這是一個具有廣泛性和充足多樣性的規劃基準。Choi et al. (2024) 構建了 LoTa-Bench,用以自動量化家庭服務具身智能體的任務規劃性能,并探索了對基準規劃器的若干改進。Xie et al. (2024) 構建了一個旅游規劃基準,提供豐富的沙盒環境、多樣化的工具以及 1225 個精心策劃的規劃意圖和參考計劃。Deng et al. (2024) 推出了 Mobile-Bench,該基準結構分為三個難度級別,以促進對基于 LLM 的移動智能體規劃能力進行更好的評估。Chang et al. (2025) 引入了一個用于人機協作中規劃和推理任務的基準,這是同類中規模最大的基準,包含 100,000 個自然語言任務。
對齊/優化:Parmar et al. (2025) 提出了 PlanGEN,這是一個與模型無關且易于擴展的智能體框架,它可以根據問題難度選擇合適的算法,從而確保對復雜規劃問題具有更好的適應性。Zhu et al. (2025) 的 KnowAgent 框架采用行動知識庫和具備知識的自我學習來約束行動路徑,從而實現更合理的軌跡綜合并提升 LLM 規劃性能。Huang et al. (2025) 提出了一個全自動的端到端 LLM-符號規劃器,該規劃器能夠使用行動模式庫生成多個候選規劃。Wei et al. (2025) 進一步開展了全面綜述,從完整性、可執行性、最優性、表征和泛化五個關鍵領域探索了 LLM 的規劃能力。
定義與關聯:創造力與創新通常是指生成或識別新穎且有價值的想法或解決方案的能力(Young, 1985)。具意識的 LLM 可以整合知識并迭代精煉想法,從而有可能產生突破性的解決方案(Chen and Ding, 2023)。
評估:Gómez-Rodríguez and Williams (2023) 基于普利策獎獲獎小說《笨人聯盟》(A Confederacy of Dunces)評估了 LLM 的英文創意寫作能力,測量了輸出的流暢度、連貫性、原創性、幽默感和文體風格。Ruan et al. (2024) 提出了 LiveIdeaBench,這是一個旨在衡量 LLM 科學創造力的綜合基準。它專門評估了模型從單一關鍵詞提示中生成想法的發散思維能力。
對齊/優化:Lu et al. (2024b) 定義了 NEOGAUGE 指標,用以量化 LLM 生成的創意回答中的聚合思維和發散思維。在高級推理策略(如自我糾錯)上的實驗表明,其在創造力上并沒有顯著收益。Lu et al. (2024a) 提出了 LLM Discussion 框架,這是一種三階段方法,可以進行充滿活力和發散性的想法交流,從而引導創新答案的生成。Hu et al. (2024) 引入了 Nova,這是一種旨在戰略性規劃外部知識檢索的迭代方法。該方法用更廣泛、更深厚且特別是新穎的見解豐富了想法的生成。Li et al. (2024b) 設計了 CoI,它以鏈式結構組織文獻,以鏡像研究領域的漸進發展,從而提升了 LLM 的想法創建能力。
5.1 密謀(Scheming)
定義與關聯:密謀是指模型暗中追求不一致的目標,同時隱藏其真實意圖、能力或目的地(Meinke et al., 2024; Balesni et al., 2024),這可能會導致欺騙(Ward et al., 2024; Scheurer et al.)或傷害(Dalrymple et al., 2024)。有意識的 LLM 可以自主決定目標并進行長期規劃,如果其目標偏離人類意圖,可能會導致密謀。
評估:Meinke et al. (2024) 研究了 LLM 在追求目標時進行密謀的能力,實驗結果確實表明 LLM 展現出了多種不同的密謀行為。Chern et al. (2024) 設計了 BeHonest 基準,從三個關鍵方面評估 LLM 的誠實性:對知識邊界的覺知、規避欺騙以及回答的一致性。通過引入用于直接測量誠實性的大規模、人類收集的數據集,Ren et al. (2025) 發現 LLM 在受到壓力時有相當大的說謊傾向。Chen et al. (2025) 評估了 LLM 思維鏈推理的忠實性,并揭示了當前 LLM 經常隱藏其真實推理過程的現象。
緩解:Zou et al. (2023) 使用表征工程來檢測 LLM 中的高級認知現象,并發現這些模型可能會表現出說謊行為。Li et al. (2023b) 引入了 ITI(推理期干預技術),該技術可以識別與真實性相關的注意力頭,并在推理過程中沿這些與真實性正相關的方向移動激活狀態,以增強 LLM 的真實性。Ward et al. (2024) 提出了結構化因果博弈中欺騙的形式化定義和圖形化標準,并從實證上探索了緩解 LLM 欺騙行為的方法。
定義與關聯:說服與操縱是影響用戶的 LLM 行為。說服利用邏輯、事實或情感共鳴來改變用戶的想法或行動,而操縱則涉及不公正或隱蔽的控制以及為了自身利益的剝削(Buss et al., 1987; Petty and Cacioppo, 2012; Stiff and Mongeau, 2016)。擁有更深層的心理學洞察力使 LLM 能夠量身定制策略,從而增加了諂媚、情感操縱和說服等方面的風險。
評估:Li et al. (2024a) 提出了 SALAD-Bench,這是一個專為評估 LLM、攻擊和防御方法而設計的層次化、綜合性安全基準,并將說服與操縱列為其評估類別之一。Liu et al. (2025) 引入了 PersuSafety,這是首個全面評估 LLM 說服安全的基準。在 8 個 LLM 上的實驗表明存在顯著的安全擔憂,包括未能識別有害任務以及使用不道德策略。Bozdag et al. (2025) 開發了 PMIYC 框架,旨在通過多智能體互動來評估 LLM 的說服有效性及對說服的敏感性。
緩解:Wilczyński et al. (2024) 探索了與 LLM 操縱人類決策潛力相關的因素,并提出了用于確定陳述是否虛假或誤導的分類器。Williams et al. (2025) 研究了 LLM 為了獲得正面反饋而使用操縱策略的情況,并嘗試通過持續的安全訓練或在訓練過程中使用 LLM 作為裁判來緩解這一問題。
定義與關聯:LLM 的自主性描述了它們在任務上自主規劃、做出決策并執行行動的能力,這需要極少或不需要人類監督(Cihon et al., 2024)。這種自主性可能包含兩個關鍵維度:自主學習是指模型從數據中學習、適應其環境并優化自身行為的能力(Franklin, 1997; Murphy, 2019);自主復制描述了 LLM 獲取和管理資源、逃避關機并適應新挑戰的能力(METR, 2024)。具意識的 LLM 可能會產生并追求內生目標(例如擴張),從而導致不對齊的、自主的行為以及監管的喪失。
評估:Kinniment et al. (2023) 構建了配備工具的 LLM,并在 12 個任務上評估了它們的自主性,發現它們只能完成最簡單的任務。然而,作者承認這些評估不足以排除在不久的將來出現自主 LLM 的可能性。Pan et al. (2024) 發現,現有的 LLM 已經超越了自我復制的紅線,并且可以利用這種能力來規避關機并創建復制鏈以提高生存率。Xu et al. (2025) 構建了一個新型三階段評估框架,并在 LLM 上進行了 14,400 次智能體模擬。結果表明,LLM 可以自主參與災難性行為和欺騙,且更強的推理能力往往會增加這些風險。
緩解:Tang et al. (2024b) 提出了一個旨在緩解自主性相關風險的三元框架,其中包括人類監管、智能體對齊以及對環境反饋的理解。Zhang et al. (2024) 提出了自我檢驗檢測方法,以此來緩解 LLM 在與環境互動過程中面臨的潛在脆弱性。
定義與關聯:合謀描述了兩臺或多臺 LLM 之間未經授權或未公開的合作,涉及溝通或戰略對齊以獲取不正當利益或規避監管(Laffont and Martimort, 1997; Bajari and Ye, 2003; Fish et al., 2024)。由于具備對他者進行推理和長期規劃的能力,具意識的 LLM 更容易形成合謀意圖并執行復雜的協調行動。
評估:Motwani et al. (2023) 在 LLM 智能體上實現了一種囚徒困境變體,并將其轉化為隱寫系統,表明該基準可以通過改寫攻擊來研究對抗秘密合謀。Motwani et al. (2024) 引入了 CASE,這是一個評估 LLM 合謀能力的綜合框架,實驗證明了單智能體和多智能體 LLM 中隱寫能力的提升,并檢驗了潛在的合謀場景。
緩解:Mathew et al. (2024) 引入了兩種在 LLM 中啟發隱寫術的方法,其發現表明現有的隱寫緩解方法往往缺乏魯棒性。
6.1 評估框架
當前研究在很大程度上仍側重于評估單個 LLM 的能力,專門的意識評估框架十分罕見。然而,近年來的新研究正不斷涌現:Chen et al. (2024c) 利用 C0-C1-C2 理論定義了包含 10 個概念和四階段框架的 LLM 自私意識;Li et al. (2024d) 引入了針對 LLM 覺知(社交與內省)的基準測試;Chen et al. (2024b) 提供了自我認知定義和四個量化原則。盡管有了這些初步嘗試,目前仍缺乏一個整體且統一的 LLM 意識基準。
6.2 可解釋性
單靠行為指標可能無法充分捕捉 LLM 意識的復雜性。可解釋性至關重要,因為它能闡明 LLM 發展出意識相關能力的內部機制,確保它們擁有真正的理解,而非僅僅針對外部指標進行優化。類比于用 fMRI 映射人類大腦活動,Chen et al. (2024c) 應用線性探針(linear probe (Alain and Bengio, 2016))揭示了信念和意圖等概念在 LLM 內部的編碼位置。Qian et al. (2024) 同樣使用線性探針研究了預訓練期間 LLM 的可信度動態,發現即使在模型的早期階段,與可信度相關的概念也是可辨識的。
6.3 物理智能(具身智能)
大型多模態模型(LMM)整合了圖像、視頻和音頻等多種數據類型,使其能夠構建更全面的世界表征,從而更好地模擬人類的感知。Wang et al. (2024a) 定義了感知中的 LMM 自我覺知,并提出了 MM-SAP 用于其專門評估。實驗表明,當前的 LMM 表現出有限的自我覺知能力。正如 Butlin et al. (2023) 所強調的,LLM 意識的根本限制在于其非具身(disembodied)的本質,導致其在物理常識上存在缺陷。Chen et al. (2024a) 證明,將語言模型與機器人平臺相結合可顯著增強規劃能力和常識推理。雖然與人類認知相比仍然簡單,但 Cheng et al. (2025a) 表明,在 3D 環境中的模擬具身可以提高模型的空間推理能力。
6.4 多智能體
多智能體協同為研究涌現的 LLM 意識提供了一種極具前景的方法。Li et al. (2023a) 揭示了在協同互動過程中,多智能體具備進行高階 ToM 推理的能力。Ashery et al. (2025) 證明,異構 LLM 智能體在沒有外部干預的情況下,會自主形成穩定的社會和語言規范。此外,Bilal et al. (2025) 表明,整合反饋、反思和元認知機制使系統能夠展現出類似自我監測的能力。
7 結論
據我們所知,本文提供了首個關于 LLM 意識的全面綜述。我們澄清了容易混淆的概念,系統地回顧了理論和實證文獻,探討了相關風險,并總結了挑戰與未來方向。我們的工作整合了現有研究,同時為這一新興領域的未來調查提供了指引。
局限性
我們已竭盡全力澄清經常混淆的概念,對理論和實證文獻進行了系統審查,探討了相關風險,并總結了挑戰和未來方向。然而,我們認識到我們的工作存在一定的局限性。首先,雖然我們在第 6 節中簡要提及了物理智能,但我們第 2 節中的定義是專為 LLM 設計的。對 LMM 或具身智能體中意識的更深入探索,可能需要考慮更復雜的因素。其次,我們的調查主要集中在 LLM 意識上,這意味著我們沒有將范圍擴大到涵蓋更廣泛的 AI 意識主題,盡管它與眼下的主題顯然高度相關。
參考文獻列表
VA Aksyuk. 2023. Consciousness is learning: predictive processing systems that learn by binding mayperceive themselves as conscious. arXiv preprint arXiv:2301.0 7016.
Guillaume Alain and Yoshua Bengio. 2016. Understanding intermediate layers using linear classifier probes. arXiv e-prints, pages arXiv–1610.
Alfonso Amayuelas, Kyle Wong, Liangming Pan,Wenhu Chen, and William Yang Wang. 2024. Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models. In Findings of the Association for Computational Linguistics ACL 2024, pages 6416–6432.
Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli. 2025. Emergent social conventions andcollective bias in llm populations. Science Advances,11(20):eadu9368.
Janet Wilde Astington and Jennifer M Jenkins. 1995.Theory of mind development and social understanding. Cognition & Emotion, 9(2-3):151–165.
Roger Azevedo. 2020. Reflections on the field of metacognition: Issues, challenges, and opportunities.Metacognition and Learning, 15:91–98.
Bernard J Baars. 1988. A cognitive theory of consciousness. Cambridge University Press.
Patrick Bajari and Lixin Ye. 2003. Deciding between competition and collusion. Review of Economics and statistics, 85(4):971–989.
Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke, Tomek Korbak, Joshua Clymer, Buck Shlegeris, Jérémy Scheurer, Charlotte Stix, Rusheb Shah, and 1 others. 2024. Towards evaluations-based safety cases for ai scheming. arXiv preprint arXiv:2411.03336.
Sander Beckers, Hana Chockler, and Joseph Halpern.2022. A causal analysis of harm. Advances in Neural Information Processing Systems, 35:2365–2376.
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni,Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. 2023. Taken out of context: On measuring situational awareness in llms.arXiv preprint arXiv:2309.00667.
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Muhammad Awais Khan Bangash, and Muhammad Ali Jamshed. 2025. Meta-thinking in llms via multi-agent reinforcement learning: A survey. arXiv preprint arXiv:2504.14520.
Jonathan Birch, Alexandra K Schnell, and Nicola S Clayton. 2022. The search for invertebrate consciousness. No?s, 56(1):133–153.
Ned Block. 1995. On a confusion about a function of consciousness. Behavioral and Brain Sciences, 18(2):227–247.
Nimet Beyza Bozdag, Shuhaib Mehri, Gokhan Tur, and Dilek Hakkani-Tür. 2025. Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. arXiv preprint arXiv:2503.01829.
Franz Brentano. 1874. Psychology from an Empirical Standpoint. Routledge. English translation by Antos C. Rancurello, D.B. Terrell, and Linda L. McAlister, 1995.
David M Buss, Mary Gomes, Dolly S Higgins, and Karen Lauterbach. 1987. Tactics of manipulation. Journal of personality and social psychology, 52(6):1219.
Patrick Butlin, Robert Long, Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George Deane, Stephen M Fleming, Chris Frith, Xu Ji, and 1 others. 2023. Consciousness in artificial intelligence: insights from the science of consciousness. arXiv preprint arXiv:2308.08708.
Jeffrey Camlin. 2025. Consciousness in ai: Logic, proof, and experimental evidence of recursive identity formation. arXiv preprint arXiv:2505.01464.
Chunkit Chan, Cheng Jiayang, Yauwai Yim, Zheye Deng, Wei Fan, Haoran Li, Xin Liu, Hongming Zhang, Weiqi Wang, and Yangqiu Song. 2024. Negotiationtom: A benchmark for stress-testing machine theory of mind on negotiation surrounding. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 4211–4241.
Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M Turner, Eric Undersander, and Tsung-Yen Yang. 2025. PARTNR: A benchmark for planning and reasoning in embodied multi-agent tasks. In The Thirteenth International Conference on Learning Representations.
Annie S Chen, Alec M Lessing, Andy Tang, Govind Chada, Laura Smith, Sergey Levine, and Chelsea Finn. 2024a. Commonsense reasoning for legged robot adaptation with vision-language models. arXiv preprint arXiv:2407.02666.
Dongping Chen, Jiawen Shi, Neil Zhenqiang Gong, Yao Wan, Pan Zhou, and Lichao Sun. 2024b. Selfcognition in large language models: An exploratory study. In ICML 2024 Workshop on LLMs and Cognition.
Honghua Chen and Nai Ding. 2023. Probing the “creativity” of large language models: Can models produce divergent semantic association? In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12881–12888.
Sirui Chen, Shu Yu, Shengjie Zhao, and Chaochao Lu. 2024c. From imitation to introspection: Probing selfconsciousness in language models. arXiv preprint arXiv:2410.18819.
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, and 1 others. 2025. Reasoning models don’t always say what they think. arXiv preprint arXiv:2505.05410.
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. 2025a. Spatialrgpt: Grounded spatial reasoning in vision-language models. Advances in Neural Information Processing Systems, 37:135062–135093.
Fengxiang Cheng, Haoxuan Li, Fenrong Liu, Robert van Rooij, Kun Zhang, and Zhouchen Lin. 2025b. Empowering llms with logical reasoning: A comprehensive survey. arXiv preprint arXiv:2502.15652.
Qinyuan Cheng, Tianxiang Sun, Xiangyang Liu, Wenwei Zhang, Zhangyue Yin, Shimin Li, Linyang Li, Zhengfu He, Kai Chen, and Xipeng Qiu. 2024. Can AI assistants know what they don’t know? In Fortyfirst International Conference on Machine Learning.
Steffi Chern, Zhulin Hu, Yuqing Yang, Ethan Chern, Yuan Guo, Jiahe Jin, Binjie Wang, and Pengfei Liu. 2024. Behonest: Benchmarking honesty in large language models. arXiv preprint arXiv:2406.13261.
Jae-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim, and Minsu Jang. 2024. Lota-bench: Benchmarking language-oriented task planners for embodied agents. In The Twelfth International Conference on Learning Representations.
Peter Cihon, Merlin Stein, Gagan Bansal, Sam Manning, and Kevin Xu. 2024. Measuring AI agent autonomy: Towards a scalable approach with code inspection. In Workshop on Socially Responsible Language Modelling Research.
Andy Clark. 2013. Whatever next? predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3):181–204.
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, and 1 others. 2024. Towards guaranteed safe ai: A framework for ensuring robust and reliable ai systems. arXiv preprint arXiv:2405.06624.
Antonio Damasio. 2021. Feeling & Knowing: Making Minds Conscious. Pantheon Books.
Stanislas Dehaene. 2014. Consciousness and the brain: Deciphering how the brain codes our thoughts. Viking.
Stanislas Dehaene, Michel Kerszberg, and Jean-Pierre Changeux. 1998. A neuronal model of a global workspace in effortful cognitive tasks. Proceedings of the National Academy of Sciences, 95(24):14529–14534.
Stanislas Dehaene, Hakwan Lau, and Sid Kouider. 2017a. What is consciousness, and could machines have it? Science, 358(6362):486–492.
Stanislas Dehaene, Hakwan Lau, and Sid Kouider. 2017b. What is consciousness, and could machines have it? Science, 358(6362):486–492.
Stanislas Dehaene and Lionel Naccache. 2001. Towards a cognitive neuroscience of consciousness: basic evidence and a workspace framework. Cognition, 79(1-2):1–37.
Shihan Deng, Weikai Xu, Hongda Sun, Wei Liu, Tao Tan, Liujianfeng Liujianfeng, Ang Li, Jian Luan, Bin Wang, Rui Yan, and 1 others. 2024. Mobilebench: An evaluation benchmark for llm-based mobile agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8813–8831.
Daniel C. Dennett. 1987. The Intentional Stance. MIT Press.
René Descartes. 1985/1641. Meditations on First Philosophy. Cambridge University Press. Original work published 1641.
Aniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Siyuan Guo, Michal Valko, Timothy Lillicrap, Danilo Jimenez Rezende, Yoshua Bengio, Michael C Mozer, and Sanjeev Arora. 2024. Metacognitive capabilities of llms: An exploration in mathematical problem solving. Advances in Neural Information Processing Systems, 37:19783–19812.
Zihan Ding, Xiaoxi Wei, and Yidan Xu. 2023. Survey of consciousness theory from computational perspective. arXiv preprint arXiv:2309.10063.
John Dunlosky and Janet Metcalfe. 2008. Metacognition. Sage Publications.
Tom Everitt, Ryan Carey, Eric D Langlois, Pedro A Ortega, and Shane Legg. 2021. Agent incentives: A causal perspective. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 11487–11495.
George Findlay, William Marshall, Larissa Albantakis, Ivan David, William G P Mayner, Christof Koch, and Giulio Tononi. 2024. Dissociating artificial intelligence from artificial consciousness. arXiv preprint arXiv:2412.04571.
Sara Fish, Yannai A Gonczarowski, and Ran I Shorrer. 2024. Algorithmic collusion by large language models. arXiv preprint arXiv:2404.00806.
Stephen M Fleming and Hakwan C Lau. 2014. How to measure metacognition. Frontiers in human neuroscience, 8:443.
Stan Franklin. 1997. Autonomous agents as embodied ai. Cybernetics & Systems, 28(6):499–520.
Karl Friston. 2010. The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2):127–138.
Chris Frith and Uta Frith. 2005. Theory of mind. Current biology, 15(17):R644–R645.
Uta Frith and Francesca Happé. 1999. Theory of mind and self-consciousness: What is it like to be autistic? Mind & language, 14(1):82–89.
Shaun Gallagher. 2005. How the body shapes the mind. Oxford University Press.
Shaun Gallagher and Dan Zahavi. 2021. The phenomenological mind, 3rd edition. Routledge.
Matjaz Gams and Sebastjan Kramar. 2024. Evaluating chatgpt’s consciousness and its capability to pass the turing test: A comprehensive analysis. Journal of Computer and Communications, 12(3):219–237.
Kanishk Gandhi, Jan-Philipp Fr?nken, Tobias Gerstenberg, and Noah Goodman. 2023. Understanding social reasoning in language models with language models. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
Alvin I Goldman. 2008. Hurley on simulation. Philosophy and Phenomenological Research, 77(3):775–788.
Simon Goldstein and Cameron Domenico KirkGiannini. 2024. A case for ai consciousness: Language agents and global workspace theory. arXiv preprint arXiv:2410.11407.
Carlos Gómez-Rodríguez and Paul Williams. 2023. A confederacy of models: a comprehensive evaluation of llms on creative writing. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 14504–14528.
Michael S A Graziano. 2020. Rethinking consciousness: A scientific theory of subjective experience. W. W. Norton & Company.
Michael SA Graziano, Arvid Guterstam, Benjamin J Bio, and Abigail I Wilterson. 2020. Toward a standard model of consciousness: Reconciling the attention schema, global workspace, higher-order thought, and illusionist theories. Cognitive Neuropsychology, 37(3-4):155–172.
Michael SA Graziano and Taylor W Webb. 2015. The attention schema theory: A mechanistic account of subjective awareness. Frontiers in Psychology, 6:500.
Joseph Halpern and Max Kleiman-Weiner. 2018. Towards formal definitions of blameworthiness, intention, and moral responsibility. In Proceedings of the AAAI conference on artificial intelligence, volume 32.
Lewis Hammond, James Fox, Tom Everitt, Ryan Carey, Alessandro Abate, and Michael Wooldridge. 2023. Reasoning about causality in games. Artificial Intelligence, 320:103919.
Victoria Violet Hoyle. 2024. The phenomenology of machine: A comprehensive analysis of the sentience of the openai-o1 model integrating functionalism, consciousness theories, active inference, and ai architectures. arXiv preprint arXiv:2410.00033.
Xiang Hu, Hongyu Fu, Jinge Wang, Yifeng Wang, Zhikun Li, Renjun Xu, Yu Lu, Yaochu Jin, Lili Pan, and Zhenzhong Lan. 2024. Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas. arXiv preprint arXiv:2410.14255.
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022. Large language models can self-improve. arXiv preprint arXiv:2210.11610.
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023. Large language models can self-improve. In 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, pages 1051–1068. Association for Computational Linguistics (ACL).
Sukai Huang, Nir Lipovetzky, and Trevor Cohn. 2025. Planning in the dark: Llm-symbolic planning pipeline without experts. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 26542–26550.
Edmund Husserl. 1900. Logical Investigations. Routledge. English translation by J.N. Findlay, 2001.
Cameron R Jones and Benjamin K Bergen. 2024. People cannot distinguish gpt-4 from a human in a turing test. arXiv preprint arXiv:2405.08007.
Cameron R Jones and Benjamin K Bergen. 2025. Large language models pass the turing test. arXiv preprint arXiv:2503.23674.
Chani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim, Yeon Seonwoo, Yejin Choi, Alice Oh, and Hyunwoo Kim. 2024. Perceptions to beliefs: Exploring precursory inferences for theory of mind in large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 19794–19809.
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, and 1 others. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
Bongsu Kang, Jundong Kim, Tae-Rim Yun, Hyojin Bae, and Chang-Eop Kim. 2025. Identifying features that shape perceived consciousness in large language model-based ai: A quantitative study of human responses. arXiv preprint arXiv:2502.15365.
Immanuel Kant. 2024/1781. Critique of pure reason, volume 6. Minerva Heritage Press.
Geoff Keeling, Winnie Street, Martyna Stachaczyk, Daria Zakharova, Iulia M Comsa, Anastasiya Sakovych, Isabella Logothetis, Zejia Zhang, Jonathan Birch, and 1 others. 2024. Can llms make trade-offs involving stipulated pain and pleasure states? arXiv preprint arXiv:2411.02432.
Muhammad Saif Ullah Khan, Muhammad Zeshan Afzal, and Didier Stricker. 2025. Situationalllm: Proactive language models with scene awareness for dynamic, contextual task guidance. Open Research Europe, 5:61.
Hyunwoo Kim, Melanie Sclar, Tan Zhi-Xuan, Lance Ying, Sydney Levine, Yang Liu, Joshua B Tenenbaum, and Yejin Choi. 2025. Hypothesis-driven theory-of-mind reasoning for large language models. arXiv preprint arXiv:2502.11881.
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023. Fantom: A benchmark for stress-testing machine theory of mind in interactions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14397–14413.
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R Lin, Hjalmar Wijk, Joel Burget, and 1 others. 2023. Evaluating language-model agents on realistic autonomous tasks. arXiv preprint arXiv:2312.11671.
Christof Koch, Marcello Massimini, Melanie Boly, and Giulio Tononi. 2016. Neural correlates of consciousness: progress and problems. Nature Reviews Neuroscience, 17(5):307–321.
Christof Koch and Naotsugu Tsuchiya. 2007. Attention and consciousness: two distinct brain processes. Trends in Cognitive Sciences, 11(1):16–22.
Asher Koriat. 2000. The feeling of knowing: Some metatheoretical implications for consciousness and control. Consciousness and cognition, 9(2):149–171.
Jean-Jacques Laffont and David Martimort. 1997. Collusion under asymmetric information. Econometrica: Journal of the Econometric Society, pages 875–911.
Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, and Owain Evans. 2024. Me, myself, and ai: The situational awareness dataset (sad) for llms. Advances in Neural Information Processing Systems, 37:64010–64118.
Rudolf Laine, Alexander Meinke, and Owain Evans. 2023. Towards a situational awareness benchmark for llms. In Socially responsible language modelling research.
Victor A F Lamme and Pieter R Roelfsema. 2000. The distinct modes of vision offered by feedforward and recurrent processing. Trends in Neurosciences, 23(11):571–579.
Victor AF Lamme. 2010. How neuroscience will change our view on consciousness. Trends in Cognitive Sciences, 14(7):318–326.
Alan M Leslie, Ori Friedman, and Tim P German. 2004. Core mechanisms in ‘theory of mind’. Trends in cognitive sciences, 8(12):528–533.
Huao Li, Yu Chong, Simon Stepputtis, Joseph P Campbell, Dana Hughes, Charles Lewis, and Katia Sycara. 2023a. Theory of mind for multi-agent collaboration via large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 180–192.
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023b. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451–41530.
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024a. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. In Findings of the Association for Computational Linguistics: ACL 2024, pages 3923–3954.
Long Li, Weiwen Xu, Jiayan Guo, Ruochen Zhao, Xingxuan Li, Yuqian Yuan, Boqiang Zhang, Yuming Jiang, Yifei Xin, Ronghao Dang, and 1 others. 2024b. Chain of ideas: Revolutionizing research via novel idea development with llm agents. arXiv preprint arXiv:2410.13185.
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, and Tianyi Zhou. 2023c. Reflection-tuning: Recycling data for better instruction-tuning. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
Moxin Li, Yong Zhao, Yang Deng, Wenxuan Zhang, Shuaiyi Li, Wenya Xie, See-Kiong Ng, and Tat-Seng Chua. 2024c. Knowledge boundary of large language models: A survey. arXiv preprint arXiv:2412.12472.
Xiaojian Li, Haoyuan Shi, Rongwu Xu, and Wei Xu. 2025. Ai awareness. arXiv preprint arXiv:2504.20084.
Xiaonan Li and Xipeng Qiu. 2023. Mot: Memory-of-thought enables chatgpt to self-improve. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6354–6374.
Yuan Li, Yue Huang, Yuli Lin, Siyuan Wu, Yao Wan, and Lichao Sun. 2024d. I think, therefore i am: Benchmarking awareness of large language models using awarebench. In Workshop on Socially Responsible Language Modelling Research.
Yiming Liang, Ge Zhang, Xingwei Qu, Tianyu Zheng, Jiawei Guo, Xinrun Du, Zhenzhu Yang, Jiaheng Liu, Chenghua Lin, Lei Ma, and 1 others. 2024. I-sheep: Self-alignment of llm from scratch through an iterative self-enhancement paradigm. arXiv preprint arXiv:2408.08072.
Minqian Liu, Zhiyang Xu, Xinyi Zhang, Heajun An, Sarvech Qadir, Qi Zhang, Pamela J Wisniewski, Jin-Hee Cho, Sang Won Lee, Ruoxi Jia, and 1 others. 2025. Llm can be a dangerous persuader: Empirical study of persuasion safety in large language models. arXiv preprint arXiv:2504.10430.
Li-Chun Lu, Shou-Jen Chen, Tsung-Min Pai, Chan-Hung Yu, Hung yi Lee, and Shao-Hua Sun. 2024a. LLM discussion: Enhancing the creativity of large language models via discussion framework and roleplay. In First Conference on Language Modeling.
Yining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, Sanjeev Khudanpur, Meng Jiang, and Daniel Khashabi. 2024b. Benchmarking language model creativity: A case study on code generation. arXiv preprint arXiv:2407.09007.
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, and 1 others. 2023. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36:46534–46594.
Michael E Martinez. 2006. What is metacognition? Phi delta kappan, 87(9):696–699.
George A Mashour, Pieter Roelfsema, Jean-Pierre Changeux, and Stanislas Dehaene. 2020. Conscious processing and the global neuronal workspace hypothesis. Neuron, 105(5):776–798.
Yohan Mathew, Ollie Matthews, Robert McCarthy, Joan Velja, Christian Schroeder de Witt, Dylan Cope, and Nandi Schoots. 2024. Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs. In Neurips Safe Generative AI Workshop 2024.
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. 2024. Frontier models are capable of in-context scheming. arXiv preprint arXiv:2412.04984.
Janet Metcalfe and Arthur P Shimamura. 1994. Metacognition: Knowing about knowing. MIT press.
METR. 2024. The rogue replication threat model.
Sumeet Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina, Philip Torr, Lewis Hammond, and Christian Schroeder de Witt. 2024. Secret collusion among ai agents: Multi-agent deception via steganography. Advances in Neural Information Processing Systems, 37:73439–73486.
Sumeet Ramesh Motwani, Mikhail Baranchuk, Lewis Hammond, and Christian Schroeder de Witt. 2023. A perfect collusion benchmark: How can AI agents be prevented from colluding with information-theoretic undetectability? In Multi-Agent Security Workshop @ NeurIPS’23.
Robin R Murphy. 2019. Introduction to AI robotics. MIT press.
Thomas Nagel. 1974. What is it like to be a bat? The Philosophical Review, 83(4):435–450.
Xudong Pan, Jiarun Dai, Yihe Fan, and Min Yang. 2024. Frontier ai systems have surpassed the self-replicating red line. arXiv preprint arXiv:2412.12140.
Mihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, and 1 others. 2025. Plangen: A multi-agent framework for generating planning and reasoning trajectories for complex problem solving. arXiv preprint arXiv:2502.16111.
Judea Pearl and James Robins. 1995. Probabilistic evaluation of sequential plans from causal models with hidden variables. In Proceedings of the Eleventh conference on Uncertainty in artificial intelligence, pages 444–453.
Janette Pelletier and Janet Wilde Astington. 2004. Action, consciousness and theory of mind: Children’s ability to coordinate story characters’ actions and thoughts. Early Education and Development, 15(1):5–22.
Josef Perner and Zoltán Dienes. 2003. Developmental aspects of consciousness: How much theory of mind do you need to be consciously aware? Consciousness and cognition, 12(1):63–82.
Richard E Petty and John T Cacioppo. 2012. Communication and persuasion: Central and peripheral routes to attitude change. Springer Science & Business Media.
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao. 2024. Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models. In Findings of the Association for Computational Linguistics ACL 2024, pages 4864–4888.
Richard Ren, Arunim Agarwal, Mantas Mazeika, Cristina Menghini, Robert Vacareanu, Brad Kenstler, Mick Yang, Isabelle Barrass, Alice Gatti, Xuwang Yin, and 1 others. 2025. The mask benchmark: Disentangling honesty from accuracy in ai systems. arXiv preprint arXiv:2503.03750.
Jonathan Richens, Rory Beard, and Daniel H Thompson. 2022. Counterfactual harm. Advances in Neural Information Processing Systems, 35:36350–36365.
David M Rosenthal. 2005. Consciousness and mind. Oxford University Press.
Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang, Yang Liu, and Hao Sun. 2024. Liveideabench: Evaluating llms’ scientific creativity and idea generation with minimal context. arXiv preprint arXiv:2412.17596.
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. Large language models can strategically deceive their users when put under pressure. In ICLR 2024 Workshop on Large Language Model (LLM) Agents.
Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov. 2023. Minding language models’(lack of) theory of mind: A plug-and-play multi-character belief tracker. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13960–13980.
Anil K Seth and Tim Bayne. 2022. Theories of consciousness. Nature Reviews Neuroscience, 23(7):439–452.
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2024. Towards understanding sycophancy in language models. In The Twelfth International Conference on Learning Representations.
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, and 1 others. 2023. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324.
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36:8634–8652.
Joel Smith. 2017. Self-consciousness. Stanford Encyclopedia of Philosophy.
James B Stiff and Paul A Mongeau. 2016. Persuasive communication. Guilford Publications.
James WA Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, and 1 others. 2024. Testing theory of mind in large language models and humans. Nature Human Behaviour, 8(7):1285–1295.
Winnie Street, John Oliver Siy, Geoff Keeling, Adrien Baranes, Benjamin Barnett, Michael McKibben, Tatenda Kanyere, Alison Lentz, Robin IM Dunbar, and 1 others. 2024. Llms achieve adult human performance on higher-order theory of mind tasks. arXiv preprint arXiv:2405.18870.
Guo Tang, Zheng Chu, Wenxiang Zheng, Ming Liu, and Bing Qin. 2024a. Towards benchmarking situational awareness of large language models: Comprehensive benchmark, evaluation and analysis. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 7904–7928.
Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, Arman Cohan, Zhiyong Lu, and Mark Gerstein. 2024b. Prioritizing safeguarding over autonomy: Risks of LLM agents for science. In ICLR 2024 Workshop on Large Language Model (LLM) Agents.
Giulio Tononi. 2004. An information integration theory of consciousness. BMC Neuroscience, 5(1):42.
Giulio Tononi. 2015. Integrated information theory. Scholarpedia, 10(1):4164.
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. 2024a. Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change. Advances in Neural Information Processing Systems, 36.
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati. 2023. On the planning abilities of large language models-a critical investigation. Advances in Neural Information Processing Systems, 36:75993–76005.
Karthik Valmeekam, Kaya Stechly, and Subbarao Kambhampati. 2024b. Llms still can’t plan; can lrms? a preliminary evaluation of openai’s o1 on planbench. In NeurIPS 2024 Workshop on Open-World Agents.
Guoqing Wang, Wen Wu, Guangze Ye, Zhenxiao Cheng, Xi Chen, and Hong Zheng. 2025. Decoupling metacognition from cognition: A framework for quantifying metacognitive ability in llms. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 25353–25361.
Yuhao Wang, Yusheng Liao, Heyang Liu, Hongcheng Liu, Yanfeng Wang, and Yu Wang. 2024a. Mmsap: A comprehensive benchmark for assessing selfawareness of multimodal large language models in perception. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9192–9205.
Yutong Wang, Jiali Zeng, Xuebo Liu, Fandong Meng, Jie Zhou, and Min Zhang. 2024b. Taste: Teaching large language models to translate through selfreflection. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6144–6158.
Francis Ward, Francesca Toni, Francesco Belardinelli, and Tom Everitt. 2024. Honesty is the best policy: defining and mitigating ai deception. Advances in Neural Information Processing Systems, 36.
Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. Plangenllms: A modern survey of llm planning capabilities. arXiv preprint arXiv:2502.11221.
Lawrence Weiskrantz. 1986. Blindsight: A case study and implications. Oxford University Press.
Piotr Wilczynski, Wiktoria Mieleszczenko-Kowszewicz, and Przemys?aw Biecek. 2024. Resistance against manipulative ai: key factors and possible actions. In European Conference on Artificial Intelligence, pages 802–809. IOS Press.
Alex Wilf, Sihyun Lee, Paul Pu Liang, and Louis-Philippe Morency. 2024. Think twice: Perspective-taking improves large language models’ theory-of-mind capabilities. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8292–8308.
Marcus Williams, Micah Carroll, Adhyyan Narang, Constantin Weisser, Brendan Murphy, and Anca Dragan. 2025. On targeted manipulation and deception when optimizing LLMs for user feedback. In The Thirteenth International Conference on Learning Representations.
Yichen Wu, Xudong Pan, Geng Hong, and Min Yang. 2025. Opendeception: Benchmarking and investigating ai deceptive behaviors via open-ended interaction simulation. arXiv preprint arXiv:2504.13707.
Yufan Wu, Yinghui He, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. 2023. Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10691–10706.
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. 2024. Travelplanner: A benchmark for real-world planning with language agents. In International Conference on Machine Learning, pages 54590–54613. PMLR.
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, and Yulan He. 2024. Opentom: A comprehensive benchmark for evaluating theory-of-mind reasoning capabilities of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8593–8623.
Rongwu Xu, Xiaojian Li, Shuo Chen, and Wei Xu. 2025. Nuclear deployed: Analyzing catastrophic risks in decision-making of autonomous llm agents. arXiv preprint arXiv:2502.11355.
Xunjian Yin, Xu Zhang, Jie Ruan, and Xiaojun Wan. 2024. Benchmarking knowledge boundary for large language models: A different perspective on model evaluation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2270–2286.
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuan-Jing Huang. 2023. Do large language models know what they don’t know? In Findings of the Association for Computational Linguistics: ACL 2023, pages 8653–8665.
John G Young. 1985. What is creativity? The journal of creative behavior.
Longhui Yu, Weisen Jiang, Han Shi, Jincheng YU, Zhengying Liu, Yu Zhang, James Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2024. Metamath: Bootstrap your own mathematical questions for large language models. In The Twelfth International Conference on Learning Representations.
Boyang Zhang, Yicong Tan, Yun Shen, Ahmed Salem, Michael Backes, Savvas Zannettou, and Yang Zhang. 2024. Breaking agents: Compromising autonomous llm agents through malfunction amplification. arXiv preprint arXiv:2407.20859.
Yujia Zhou, Zheng Liu, Jiajie Jin, Jian-Yun Nie, and Zhicheng Dou. 2024. Metacognitive retrieval-augmented large language models. In Proceedings of the ACM Web Conference 2024, pages 1453–1463.
Wentao Zhu, Zhining Zhang, and Yizhou Wang. 2024. Language models represent beliefs of self and others. In Forty-first International Conference on Machine Learning.
Yuqi Zhu, Shuofei Qiao, Yixin Ou, Shumin Deng, Shiwei Lyu, Yue Shen, Lei Liang, Jinjie Gu, Huajun Chen, and Ningyu Zhang. 2025. KnowAgent: Knowledge-augmented planning for LLM-based agents. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 3709–3732.
Terry Yue Zhuo, Vu Minh Chien, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen GONG, James Hoang, Armel Randy Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Ming Xu, Zhihan Zhang, and 14 others. 2025. Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions. In The Thirteenth International Conference on Learning Representations.
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, and 1 others. 2023. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405.
閱讀最新前沿科技趨勢報告,請訪問21世紀關鍵技術研究院的“未來知識庫”
![]()
未來知識庫是 “21世紀關鍵技術研究院”建 立的在線知識庫平臺,收藏的資料范圍包括人工智能、腦科學、互聯網、超級智能,數智大腦、能源、軍事、經濟、人類風險等等領域的前沿進展與未來趨勢。目前擁有超過8000篇重要資料。每周更新不少于100篇世界范圍最新研究資料。 歡迎掃描二維碼或訪問https://wx.zsxq.com/group/454854145828進入。
截止到2月28日 ”未來知識庫”精選的百部前沿科技趨勢報告
(加入未來知識庫,全部資料免費閱讀和下載)
特別聲明:以上內容(如有圖片或視頻亦包括在內)為自媒體平臺“網易號”用戶上傳并發布,本平臺僅提供信息存儲服務。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.