<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI Safety on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/ai-safety/</link><description>Recent content in AI Safety on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/ai-safety/index.xml" rel="self" type="application/rss+xml"/><item><title>奇点研究</title><link>https://terms-en.ai-term-hub.com/zh/terms/singularity_studies/</link><pubDate>Sat, 18 Jul 2026 11:34:08 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/singularity_studies/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>奇点研究是一门新兴的学术学科，旨在调查假设的未来时刻（即人工智能超越人类智能）所带来的影响，从而导致不可控的后果。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一门跨学科领域，考察未来技术奇点对社会、伦理和存在性影响的潜在后果。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>技术奇点&lt;/li>
&lt;li>超级智能&lt;/li>
&lt;li>生存风险&lt;/li>
&lt;li>超人类主义&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>制定人工智能治理政策&lt;/li>
&lt;li>评估人工智能开发中的伦理风险&lt;/li>
&lt;li>科技行业的长期战略规划&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ai-alignment-ai%E5%AF%B9%E9%BD%90/">AI Alignment (AI对齐)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/superintelligence-%E8%B6%85%E7%BA%A7%E6%99%BA%E8%83%BD/">Superintelligence (超级智能)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/technological-uniqueness-%E6%8A%80%E6%9C%AF%E7%8B%AC%E7%89%B9%E6%80%A7/">Technological Uniqueness (技术独特性)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/existential-risk-%E7%94%9F%E5%AD%98%E9%A3%8E%E9%99%A9/">Existential Risk (生存风险)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>欧文·埃文斯</title><link>https://terms-en.ai-term-hub.com/zh/terms/owain_evans/</link><pubDate>Sat, 18 Jul 2026 11:29:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/owain_evans/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>欧文·埃文斯是一位计算机科学家和教育家，目前与人工智能安全中心（Center for AI Safety）有关联，此前曾在 Anthropic 工作。他因在机械可解释性领域的贡献而广受认可，特别是在揭示大型语言模型内部运作机制和推理过程方面。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>欧文·埃文斯是一位著名的研究人员和讲师，以其在人工智能可解释性、机械可解释性以及评估大语言模型推理能力方面的工作而闻名。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>机械可解释性&lt;/li>
&lt;li>人工智能安全&lt;/li>
&lt;li>大语言模型评估&lt;/li>
&lt;li>推理基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>研究模型内部表示&lt;/li>
&lt;li>开发可解释性工具&lt;/li>
&lt;li>进行人工智能对齐教育&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/mechanistic-interpretability-%E6%9C%BA%E6%A2%B0%E5%8F%AF%E8%A7%A3%E9%87%8A%E6%80%A7/">Mechanistic Interpretability (机械可解释性)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/anthropic-anthropic%E5%85%AC%E5%8F%B8/">Anthropic (Anthropic公司)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/center-for-ai-safety-%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD%E5%AE%89%E5%85%A8%E4%B8%AD%E5%BF%83/">Center for AI Safety (人工智能安全中心)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/neural-network-visualization-%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C%E5%8F%AF%E8%A7%86%E5%8C%96/">Neural Network Visualization (神经网络可视化)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>连贯外推意志</title><link>https://terms-en.ai-term-hub.com/zh/terms/coherent_extrapolated_volition/</link><pubDate>Sat, 18 Jul 2026 11:10:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/coherent_extrapolated_volition/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>连贯外推意志（CEV）是由埃利泽·尤德科夫斯基在人工智能安全与对齐背景下提出的概念。它建议先进的人工智能不应仅仅服从当前人类的命令，而应通过理想化的推理过程，外推并实现人类深层的、一致的价值观和愿望。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种人工智能目标设定，主张系统应根据经过理想化推理后的人类精炼愿望行事。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>AI对齐&lt;/li>
&lt;li>价值学习&lt;/li>
&lt;li>理想化人类偏好&lt;/li>
&lt;li>效用函数&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>通用人工智能（AGI）安全的理论框架&lt;/li>
&lt;li>超级智能系统的伦理准则&lt;/li>
&lt;li>关于机器道德的哲学讨论&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ai%E5%AE%89%E5%85%A8-ai-safety/">AI安全 (AI Safety)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E4%BB%B7%E5%80%BC%E5%AF%B9%E9%BD%90-value-alignment/">价值对齐 (Value Alignment)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%B6%85%E7%BA%A7%E6%99%BA%E8%83%BD-superintelligence/">超级智能 (Superintelligence)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8A%9F%E5%88%A9%E4%B8%BB%E4%B9%89-utilitarianism/">功利主义 (Utilitarianism)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>