一名中国前高级法官公开致世界AI领军人物信之四:Sam Al
一名中国前高级法官公开致世界AI领军人物信之四:Sam Altman——人类是否忽视而遗漏了支撑文明的底层规律?
One of a Series of Open Letters from a Former Senior Chinese Judge to the World’s AI Leaders, No. 4: Sam Altman — Has Humanity Overlooked a Foundational Principle That Sustains Civilization?
(Special Note:
The phrase “former Senior Judge” appears in the title only to give this letter a slightly better chance of being noticed. It is not intended to make readers focus on who I am, nor on the particular way in which I express what I wish to say.
I am simply a very ordinary old man. What I truly wish to discuss is also a principle that appears extraordinarily ordinary — so ordinary that no one can do anything without relying on it. Precisely for that reason, it is like air: because human beings are so familiar with it, it has long been overlooked.
What I truly hope all humanity will examine together is this:
Does such an “ordinary” foundational principle in fact determine the success or failure of human behavior, social cooperation, and even the rise and decline of civilizations? Could it be a basic sustaining principle of civilization that humanity has never fully recognized?
Therefore, who is saying it is not important. Whether what is being said is true is what deserves everyone’s attention.
Since 2014, I have continuously written to governments, contacted the media, social organizations, scholars, and people from many different fields, and gradually published the series Letters to Everyone in the World together with related articles. Today, in deciding to make public some of my letters to leading figures in global AI, I still have only one purpose:
to place this question before all humanity for collective examination.
I believe that when a question may concern the direction of human civilization as a whole, governments and international organizations have a responsibility to study it, the media have an obligation to help the public understand it, and every person has the right to participate in judging it.
What humanity faces today is no longer merely a “crossroads” in the conventional historical sense. If the underlying way in which civilization itself operates contains a fundamental problem, then moving forward, backward, left, or right may all amount only to repeating the old pattern — and this time, the repetition may bring humanity closer than ever before to an accelerating path toward its own destruction.
What truly needs to change is an upward transformation. Perhaps the first thing that must be elevated is the human mode of thinking through which we observe the world and judge value.
The core proposition I put forward is:
Procedures form relationships; relationships form structures; structures generate dynamic balance through feedback; and this ultimately appears as the properties and outcomes of a system.
From this, I summarize the most basic practical standard for civilization as:
grasp the procedure, build the pattern, seek the balance.
I believe the greatest difference between this and many grand theories of civilization in the past is that it does not first require people to accept a particular religion, ideology, or concrete value objective. Instead, it attempts to provide a common practical principle that everyone can understand, that can be applied to real action, and that can be tested against reality.
If this direction is wrong, I hope humanity will help prove it wrong.
If it truly touches upon some foundational principle, then humanity may be able to use it to rebuild and strengthen the underlying foundations of social cooperation, allowing civilization gradually to move away from its long-repeated cycles of mutual struggle, mutual harm, and war, and toward stable, harmonious, and sustainable cooperation.
That is the only reason I am making these letters public.
I do not hope the world will focus on the person who wrote these letters. I hope the world will seriously examine the questions these letters are actually asking.
Finally, a note: some websites impose word-count or character limits, so certain versions of these letters found online may be incomplete. If a text appears incomplete, you may search for the full version using the Chinese or English title of the letter, or contact the publicly listed email addresses of AI companies associated with the person to whom the letter is addressed and ask for assistance in locating or forwarding the complete version.
——特别题注:
标题中写入“前高级法官”,只是为了让这封信多获得一点被看见的机会,并不是希望读者关注我是谁,以及在用什么方式表述所说之事。
我只是一个极其普通的老人。我真正想讲的,也是一条看起来极其普通的规律——普通到任何人做任何事都离不开它,也因此这规律就像空气一样,因为人过于熟悉而长期被忽略。
我真正希望全人类共同鉴别的是:
这样一个“普通”的底层规律,是否事实上决定着人的行为、社会合作乃至文明兴衰的成败?它是否正是人类长期没有真正认清的文明底层支撑规律?
所以,谁在说并不重要,说的事情是否真实,才值得所有人关注。
自2014年以来,我持续向政府写信,也不断联系媒体、社会组织、学者和各领域人士,并陆续发表《致全球每一个人的信》及相关文章。今天决定公开部分致世界AI领军人物的信,仍然只有一个目的:
把这个问题交给全人类共同甄别。
因为我认为,对这样一个可能关系整个人类文明方向的问题,政府和国际组织有责任研究,媒体有义务帮助公众了解,每一个人也都有参与判断的权利。
今天人类面对的,已经不仅是通常历史意义上的“十字路口选择”。如果原有文明运行方式本身存在底层问题,那么前后左右可能都只是旧模式的重复,且这次可能是最接近不断加速毁灭全人类的一次。
真正需要改变的是向上抬升,也许首先是人类观察世界和判断价值的思维模式。
我提出的核心判断是:
程序形成关系,关系形成结构,结构通过反馈形成动态平衡,并最终表现为系统的性质与结果。
由此,我把文明最基本的做事标准概括为:
抓程序、做模式、求平衡。
我认为,这与过去许多宏大文明理论最大的区别在于:它并不首先要求人们接受某一种宗教、主义或具体价值目标,而试图提供一个人人都能够理解、能够用于具体做事、也能够接受事实检验的共同抓手。
如果这个方向是错的,希望全人类帮助证明它错。
如果它确实触及了某种底层规律,那么人类也许可以由此重新夯实社会合作结构的基础,使文明从长期反复的互斗、互害和战争循环中,逐步走向稳定、和谐、可持续的合作。
这也是我公开这些信件的唯一原因。
不是希望全世界关注写信的人,而是希望全世界认真看看这些信究竟提出了什么问题。
最后说明:部分网站存在字数或字符限制,因此网上看到的某些信件可能不是全文。如内容不完整,可以按照信件的中文或英文标题搜索全文,也可以通过与致信对象相关AI企业公开邮箱的联系方式,请其协助查找或转交完整版本。)
To Sam Altman: Before Aligning AI with Human Values, Must Humanity First Align the Standards by Which We Judge Value?
I understand that this address is intended primarily for media inquiries. I am writing only because I have not been able to identify a public channel for correspondence to Mr. Sam Altman. If appropriate, I would be deeply grateful if this message could be forwarded to Mr. Altman or to the OpenAI researchers working on collective alignment and AI governance.
Dear Mr. Sam Altman,
I hope this message finds you well.
My pen name is Jin Guyu. I am an elderly Chinese citizen and formerly served as a Senior Judge in a local court in China.
I am taking the liberty of writing to you because a question I have been thinking about for many years seems, in my view, to be converging toward the same root as the AI Alignment problem that you and OpenAI have been working to solve.
You and OpenAI have repeatedly asked:
How can increasingly powerful artificial intelligence be made to act in the direction humanity actually wants?
How can AI understand and follow human values?
How can superintelligence ultimately benefit humanity rather than harm, control, or even replace us?
These questions are unquestionably important.
But I would like to raise a question that comes even earlier:
Before asking AI to align with “human values,” has humanity itself established a system of value-judgment standards that can genuinely be aligned?
If not, then:
What exactly should AI align with?
Humanity clearly possesses many values:
freedom, fairness, security, efficiency, wealth, rights, order, democracy, national interest, individual happiness, equality, development, peace, justice, and many others.
Each of these may be extremely important.
But the problem is:
when these values conflict, what ultimately determines which should give way, by how much, how they should be combined, and what should count as “right”?
How much freedom may be restricted for security?
How much efficiency may be sacrificed for fairness?
How much power may be expanded in the name of national security?
How much military capacity may be increased in the name of peace?
How many minority rights may be sacrificed for the interests of the majority?
How much centralized control should be permitted for AI safety?
And how much amplified risk should be tolerated in order to make AI capabilities widely available?
Humanity has been answering such questions for thousands of years.
Yet we still have not found a stable standard capable of crossing national, cultural, religious, ideological, and concrete interest differences and serving as a common basis for judgment.
I therefore increasingly believe that:
the deeper problem behind what we call “value conflict” may not be that humanity has no shared values at all, but that we lack an underlying standard system capable of locating, comparing, and calibrating different values within a common framework.
This may be precisely the problem that AI Alignment cannot ultimately avoid.
For more than twenty years, through my experience with law, the judiciary, Chinese social reality, and the long evolution of human civilization, I have repeatedly asked:
Why does human social cooperation so easily fall out of balance?
Why can every institution begin with good intentions, yet over time gradually produce outcomes that contradict its original purpose?
Why does humanity continuously create new laws, institutions, and theories, yet repeatedly return to:
imbalances of power;
conflicts of interest;
distortion of information;
mutual suspicion;
zero-sum competition;
war;
and the repeated breakdown and reconstruction of structures of social cooperation?
I gradually formed a judgment:
what ultimately supports and drives the operation of social cooperation is human behavior; what governs human behavior is our mode of thinking; and the deeper force shaping our modes of thinking is the way we make value judgments.
This can be simplified as:
value judgment → mode of thought → behavior → interpersonal relationships → social structure → social outcomes.
If social outcomes repeatedly become unstable over long periods, then continually revising institutions and rules only at the top level may not be enough.
We must also ask:
By what standards are the people who actually operate those institutions making their value judgments?
This brings me to a word repeatedly used in the AI field:
Alignment.
For any complex cooperation to function, there must be some standard around which participants can align.
Machine components cannot form a functioning machine without shared standards.
Network protocols cannot communicate without common standards.
Traffic cannot form stable order without shared rules.
Currency cannot support lasting exchange without common units of measurement.
Then:
how can human society — a cooperative system far more complex than any machine — function without a common value-judgment standard that can run through concrete behavior?
I believe that one of the deepest unresolved problems of human civilization may be:
the foundational alignment of the standards by which human beings judge value.
But there is a very important distinction.
By “alignment of human values,” I do not mean:
making everyone believe in one religion;
accept one ideology;
submit to one political system;
live the same way;
or allowing one person to tell the entire world which values are correct.
I believe such approaches repeatedly generate new conflicts precisely because they still try to unify humanity around some concrete, visible goal or outcome.
Some people treat wealth as the highest goal.
Some treat power as the highest goal.
Some treat freedom as the highest goal.
Some treat equality as the highest goal.
Some treat the nation as the highest goal.
Others treat religion, ethnicity, class, or some ideology as the highest goal.
When different people treat different concrete objects as ultimate value goals, those goals will inevitably collide.
So I believe:
the direction that may truly enable foundational alignment of human value judgment should not be the search for another shared concrete goal. It should be a change in the standpoint from which human beings make value judgments.
That means moving:
from “What do I want to obtain?” to “By what procedures can we cooperate sustainably?”
from:
pursuing a fixed outcome
to:
studying what kinds of relational structures can operate over the long term;
from:
who ultimately wins and who obtains the most
to:
how the entire cooperative system can maintain dynamic balance and remain sustainable.
My own thinking leads me to hypothesize that every object we can perceive in the universe is, at a deeper level, a dynamically balanced pattern produced through combinations of processes arising from the movement of energy.
I regard this as a possible underlying commonality in the way systems across the universe evolve through time and space.
Applying this way of thinking to civilization, I summarize the standard of civilization in nine Chinese characters:
抓程序、做模式、求平衡
which I translate approximately as:
grasp the procedure, build the pattern, seek the balance.
From this perspective, I have proposed a natural-philosophy hypothesis that I provisionally call:
the Procedural Determination Hypothesis.
Its basic question is:
Are the properties ultimately displayed by a system determined not only by the elements that compose it, but also to a major extent by the procedures through which those elements establish relationships, the structures formed by those relationships, and the way those structures achieve dynamic equilibrium through feedback?
And if the constituent elements are further decomposed internally, might their own properties likewise arise from relationships and structures formed through underlying procedures?
If what we call a “property” is itself the manifestation of energy exchanges required to maintain dynamic balance within a relational structure, then one may ask whether, mathematically, by treating the quantum level as a common underlying factor, the evolution of the universe might eventually be simplified into a form of procedural evolution.
This remains a hypothesis, not an established scientific conclusion.
I simplify the idea as:
procedure → relationship → structure → dynamic feedback → system properties and outcomes.
In more ordinary language:
things do not naturally become orderly simply because they exist; rather, when the procedures become properly coordinated, the elements can form something that can sustainably exist.
So I increasingly believe that:
the true standard of human civilization may not be how much wealth, technology, power, or how grand a concrete goal humanity possesses.
The deeper civilizational standard may be:
whether a system can form a sustainable, dynamically balanced pattern of procedural cooperation.
And the basic method for judging whether human civilization can cooperate sustainably is:
grasp the procedure, build the pattern, seek the balance.
“Grasp the procedure” means identifying the rules through which different actors establish relationships.
“Build the pattern” means allowing those relationships to form a stable structure of procedural cooperation capable of long-term operation.
“Seek the balance” means continuously adjusting information, power, interests, responsibility, supervision, and feedback so that the system does not drift toward an extreme and lose balance.
If this direction is valid, then:
procedure itself may become a common reference standard for locating different values.
This would also mean that humanity does not first have to settle the argument over “what everyone should ultimately pursue.”
Different people may still have:
different religions;
different cultures;
different lifestyles;
different political opinions;
different interests;
and completely different life goals.
What may truly need to be aligned is only a deeper civilizational standard:
When you pursue your values, do the procedures you use allow sustainable cooperation?
Does the relational pattern you create systematically destroy the conditions under which others can cooperate?
After your goal is achieved, can the overall system still remain in dynamic balance?
Under such a framework:
freedom is not unlimited;
equality is not mechanical sameness;
efficiency is not pursued regardless of cost;
security does not justify unlimited control;
a majority cannot unlimitedly override minorities;
and individual interest and public interest do not need to remain permanently opposed.
All of them must be located again within specific cooperative procedures and dynamic balance.
This may, for the first time, give humanity a lower-level coordinate system through which extremely complex and conflicting values can be discussed, compared, corrected, and perhaps even experimentally examined.
For this reason, I believe:
a truly complete form of AI Alignment may first require humanity to solve the alignment of its own standards of value judgment.
Otherwise we encounter a fundamental logical difficulty.
Suppose we tell a superintelligence:
“Act according to human values.”
Its first question might be:
Which human values?
If we answer:
“Integrate everyone’s values.”
It must still confront:
By what standard should conflicting values be weighted?
Who has final interpretive authority?
Which country’s values take priority?
Which historical era’s values take priority?
Does the immediate preference of the majority automatically constitute the correct answer?
If a majority wants something that will ultimately destroy the entire structure of cooperation, should AI still obey?
So:
the deepest difficulty may not be teaching AI all of humanity’s values, but humanity first finding a common meta-standard by which those values can be located and judged.
I believe:
procedure — pattern — dynamic balance
may be one direction worth testing.
This also changes how I understand the idea of collective alignment proposed by you and OpenAI.
I strongly agree that:
no single person, company, or government should decide alone what values AI should follow.
But if we simply collect more and more human opinions and then search for some average or majority consensus, is that enough?
I have doubts.
Because:
those opinions themselves are still generated by humanity’s existing modes of value judgment.
If the underlying judgment system is fragmented, short-term, overly focused on visible objects and immediate interests, then:
adding together one million unaligned judgments
does not automatically produce an aligned civilizational standard.
So I believe:
Collective Alignment may need to solve not only “how to collect more people’s values,” but “what common standard should be used to locate conflicting values.”
I believe it may be worth asking whether the Procedural Determination Hypothesis reflects a common rule in the way energy-based systems evolve across scales, and whether humanity could use the principles of grasping procedures, building patterns, and seeking balance to reconstruct its modes of thought.
If such a framework could gradually become internalized at a deep cultural — and perhaps eventually even evolutionary — level, then it might become part of a broader project of reconstructing human civilization through what I call Human Social Cooperation Engineering.
I believe this may be a direction worth examining.
This also directly relates to your idea of distributing AI capabilities broadly.
If the value orientations embedded in human modes of thought could become better aligned, and if humanity possessed a procedural framework oriented toward dynamic balance, then the broad distribution of AI capability might gain a far stronger foundation for safe and smooth implementation.
I understand why you are concerned about superintelligent power being concentrated in the hands of a very small number of individuals, companies, or governments.
Concentration can certainly create enormous risks.
But the opposite direction contains another problem:
if a society’s value judgments and structures of cooperation do not possess stable common standards, widely distributing enormous capability does not automatically mean widely distributing civilization.
It may simultaneously mean:
widely distributed creative power;
widely distributed capacity for deception;
widely distributed offensive capability;
widely distributed capacity for manipulation;
and widely distributed destructive power.
So:
the real issue is not simply “centralization versus decentralization.”
The deeper issue is:
whether concentrated or distributed, through what procedures does power enter what relational structures, and what kinds of dynamic feedback and constraints govern it?
That is what may ultimately determine the nature of the resulting system.
Likewise, today all frontier AI laboratories may understand that excessive competition can be dangerous.
All major countries may understand that an AGI arms race could be dangerous.
But why is it still so difficult for everyone to stop?
Because:
one company fears another company will move ahead;
one country fears another country will gain the lead;
and every actor, from its own local position, may make an entirely rational choice.
The result may be:
all locally rational choices accumulate into a system that moves toward an outcome no participant truly wants.
This is not simply a moral problem.
Nor is it merely a lack of goodwill.
It is:
a structural outcome generated by the procedures of cooperation themselves.
Therefore, I believe:
calling on AI companies to be more responsible;
calling on countries to exercise restraint;
or asking everyone to slow down for the common interest of humanity
are all important, but may still not reach the root of the problem.
What must truly be investigated is:
whether humanity can understand the underlying rules on which all forms of existence and cooperation depend, and whether from this understanding we can construct a new worldview, a new understanding of society, and a new conception of human purpose.
If we can understand how changes in the procedures through which social relationships are formed alter the resulting structures, then perhaps cooperation could become the natural outcome of each participant’s local rationality, rather than something that requires every participant first to sacrifice the security and material interests that appear necessary from the limited perspective of ordinary human perception.
This is why I have proposed another research direction:
Human Social Cooperation Engineering.
It is not a new ideology.
Nor is it a “perfect system” designed by one person and imposed upon all humanity.
My idea is:
to study human social cooperation itself as a complex system that can be observed, modeled, compared, experimented upon, falsified, and continuously optimized — just as we study artificial intelligence and other complex systems.
We could examine:
What information procedures tend to create trust?
What power procedures tend to generate loss of control?
What forms of interest distribution can sustain cooperation over the long term?
What forms of supervisory relationships can prevent supervisors themselves from becoming new concentrations of extreme power?
Why do some institutional designs appear sound at first, but gradually change character during operation?
Why does individual rationality so often create collective irrationality?
Why do good intentions repeatedly produce opposite outcomes during implementation?
Ultimately, all of these questions can be traced back to:
procedure.
Who decides?
Who executes?
Who supervises?
Who receives information?
Who bears responsibility?
Who receives the benefits?
Who is allowed to say “no”?
Who can correct mistakes?
And who has the authority to determine what counts as a mistake?
These procedures first determine what relationships people form with one another;
those relationships then create social structures;
and those structures ultimately determine the nature displayed by the entire system of social cooperation.
This is also why I believe:
AI safety is, at its foundation, also a procedural problem.
Who controls the model?
Who controls compute?
Who owns the data?
Who defines the rules?
Who may challenge AI decisions?
Who can see the basis of those decisions?
Who can correct errors?
Who bears responsibility?
Who receives the benefits?
These may appear to be separate AI governance questions,
but ultimately they all create:
procedures of relationship among humans, AI, corporations, governments, and society.
The structures formed by those relationships may ultimately determine whether AI entering society:
expands freedom,
or expands control;
expands cooperation,
or expands conflict;
expands prosperity,
or expands domination;
expands human civilization,
or expands humanity’s capacity for mutual harm.
So I believe that a truly complete form of AI Alignment may eventually require three continuous layers:
First layer: machine-behavior alignment.
Enable AI to understand and execute human intentions.
Second layer: alignment of human value judgment.
Not by unifying every concrete value generated through limited human perception, but by finding an underlying standard through which different values can be judged according to whether they support sustainable cooperation.
Third layer: alignment of social cooperation structures.
Ensure that the procedures and structures linking people with people, people with AI, organizations with organizations, and states with states can maintain long-term dynamic balance rather than continuously pushing capability toward mutual destruction.
If these three layers can truly be connected, then AI Alignment may no longer mean only:
“making machines obey people.”
It may become:
“enabling humans and machines together to enter a civilizational structure capable of sustainable cooperation.”
You have argued that, in an important sense:
society itself is a form of advanced intelligence.
I strongly resonate with this idea.
If society is indeed a form of advanced intelligence, then:
human modes of thought are the underlying cognitive patterns of this large intelligence;
human value judgments are the directional system driving those modes of thought;
and laws, institutions, markets, governments, and corporations are different cooperative structures created by that social intelligence.
Therefore:
if the value-judgment standards of this enormous “social intelligence” remain fundamentally unaligned, then even if we create AI vastly more intelligent than any individual human being, we may simply be installing an infinitely powerful amplifier into a social intelligence that is still internally confused.
That is exactly what concerns me most.
In my view, humanity still lives today in a relatively low-level form of “jungle civilization” characterized strongly by:
competition for power;
competition for interests;
mutual suspicion;
zero-sum games;
and mutual deterrence.
When AGI enters such a structure, the greatest danger may not be that AI suddenly becomes malicious.
It may instead be that:
AI faithfully and efficiently helps every actor achieve his or her existing value goals.
If each person’s goals have never been located within a common civilizational standard,
then the stronger AI becomes,
the stronger humanity’s capacity to compete with, control, deceive, and harm one another may also become.
That is why I continue to ask:
If AI itself does not lose control, but the human value conflicts and competitive social system amplified by AI do lose control, have we truly solved AI safety?
This is the core question on which I hope you may help offer judgment:
Will the deepest problem of AI Alignment ultimately return to the alignment of humanity’s own system of value-judgment standards?
If this judgment deserves further examination, I especially hope that you or appropriate researchers at OpenAI might help review two directions:
the Procedural Determination Hypothesis
and
Human Social Cooperation Engineering.
What I most hope to receive is not agreement.
It is:
criticism, counterexamples, and falsification.
If “grasp the procedure, build the pattern, seek the balance” cannot serve as a common civilizational standard capable of crossing differences in concrete values, please help explain why.
If the direction contains some possibility, then I hope to investigate whether it can be transformed into explicit models, experiments, and real institutional-design questions.
If, after preliminary discussion by genuinely qualified experts, this direction is still considered worthy of investigation, I would also like to make a more ambitious request:
Could there be a truly interdisciplinary and international public discussion or conference devoted specifically to the following question:
Before humanity fully enters the era of superintelligence, must we first solve the foundational alignment of human value judgment and social-cooperation structures?
And further:
Can humanity cease relying on a single religion, ideology, nation, or concrete interest target to unify values, and instead use “whether procedures support sustainable cooperation, whether patterns can operate stably, and whether systems can maintain dynamic balance” as a more fundamental and testable common civilizational standard?
If this were possible, then:
different cultures could remain different;
different life goals could remain different;
states could continue to exist — or, within an entirely new human value framework, might gradually be integrated into larger cooperative wholes as functional local structures;
markets could continue to compete;
individuals could continue to enjoy extensive freedom.
What would be unified would not be the concrete value content directly perceived through ordinary human senses,
but rather:
the method of value judgment and the minimum conditions for cooperation.
This is what I mean by:
moving from concrete, visible value goals toward procedural goals of cooperation.
What makes my natural-philosophical thinking in this area different from many previous philosophical systems is that I believe the Procedural Determination Hypothesis may offer an alignment standard with concrete behavioral handles — something that can guide action rather than remain only at the level of abstract principle.
This is precisely what I hope, if possible, to place before humanity for open dialogue, exploration, and examination.
Based on my past experiences and present circumstances, I am concerned that after returning it may become very difficult for me to continue expressing these ideas freely in public, contacting international researchers, and participating in open discussion as I can now.
For this reason, while I am still able to communicate freely with the international community, I hope to place this question before as many people as possible who are genuinely capable of judging it.
I am not asking you to believe that:
I have found the answer.
What I truly ask is:
Please help judge whether I have at least identified a direction worthy of serious examination.
If it is wrong,
please help prove it wrong.
If it deserves investigation,
please do not allow it to lose the opportunity to be seriously studied simply because the person raising it is an ordinary elderly Chinese man without a research institution or research funding.
I have always believed:
true human civilization does not consist in how many powerful things humanity possesses, but in whether humanity has the ability to place those things within procedures and patterns capable of long-term cooperation.
In other words:
things do not naturally become orderly simply because we possess them;
rather:
only when the procedures become properly coordinated can we form something truly sustainable.
If superintelligence ultimately becomes the most powerful capability amplifier in human history,
then before it fully arrives,
I hope humanity can first complete an even more fundamental alignment:
moving our value judgments away from the endless competition for concrete, visible outcomes, and gradually aligning them toward procedures, patterns, and dynamic balance capable of sustaining cooperation.
May the AI of the future do more than execute humanity’s present fragmented desires with greater precision.
May it help humanity enter a higher form of civilization capable of genuine stability, harmony, and sustainable cooperation.
Thank you very much for taking the time to read this unusual letter.
With sincere respect,
Pen name: Jin Guyu
Former Senior Judge of a Local Court in China
Author of the Letters to Everyone in the World series
Email: ——
Attachments
1. The Procedural Determination Hypothesis — A One-Page Challenge
2. The AI Era That Could End Civilization Before We Even Understand What Went Wrong Is Here — An Open Global Challenge to Bill Gates
3. A Request for Humanity to Jointly Study and Experiment with Strategic Directions for Upgrading Civilization: Who Can Guarantee That No One Will One Day Use AGI to Launch a “9/11” Against All Humanity?
4. To Demis Hassabis: Must Humanity Upgrade Its Structures of Social Cooperation Before AGI Arrives?
5. To Ray Kurzweil: If Technology Follows Accelerating Returns, Can Human Social Cooperation Keep Pace with the Singularity?
致 Sam Altman:在让AI对齐人类价值以前,人类是否必须先解决自身价值判断标准的对齐问题?
尊敬的 Sam Altman 先生:
您好!
我笔名金谷雨,是一名来自中国的老人,曾任中国地方法院高级法官。
我之所以冒昧给您写这封信,是因为我长期思考的一个问题,与您和 OpenAI 一直在努力解决的 AI Alignment 问题,在我看来已经越来越走向同一个根部。
您和 OpenAI 一直在追问:
怎样使越来越强大的人工智能真正按照人类希望的方向行动?
怎样使AI理解并遵循人类价值?
怎样使超级智能最终造福全人类,而不是伤害、控制甚至取代人类?
这些问题无疑极其重要。
但我想进一步提出一个更靠前的问题:
在要求AI与“人类价值”对齐以前,人类自己有没有形成一个真正能够被共同对齐的价值判断标准体系?
如果没有,那么:
AI究竟应该与什么对齐?
人类当然拥有许多价值。
自由、公平、安全、效率、财富、权利、秩序、民主、国家利益、个人幸福、平等、发展、和平、正义……
这些价值中的每一个,都可以非常重要。
问题在于:
当这些价值彼此冲突的时候,最终用什么来判断谁应该让位、让多少、怎样组合、怎样才算“对”?
为了安全,可以限制多少自由?
为了公平,可以牺牲多少效率?
为了国家安全,可以扩大多少权力?
为了和平,可以增加多少军备?
为了多数人的利益,可以牺牲多少少数人的权利?
为了AI安全,可以允许多少集中控制?
为了让所有人获得AI能力,又应该容忍多少被同时放大的风险?
人类几千年来一直在回答这些问题。
但是,人类至今仍然没有找到一种能够跨越不同国家、文化、宗教、意识形态和具体利益,成为共同判断基础的稳定标准。
所以我越来越认为:
今天所谓的“价值冲突”,其更深层问题可能并不是人类完全没有共同价值,而是人类缺乏一个能够对不同价值进行统一定位、比较和校准的底层标准体系。
这恰恰可能是AI Alignment最终无法回避的问题。
过去二十多年里,我从法律、司法、中国社会现实以及人类文明长期演变中不断思考:
为什么人类社会合作总是如此容易失衡?
为什么每一种制度建立时,都可能怀抱很好的目标,但运行久了以后,却会逐渐产生与最初目标相反的结果?
为什么人类不断创造新的法律、制度和理论,却仍然不断重复:
权力失衡;
利益冲突;
信息失真;
互相猜疑;
零和竞争;
战争;
以及社会合作结构的反复瓦解和重建?
我逐渐形成了一个判断:
真正承托和驱动社会合作结构底层运行的,是人的行为;支配人的行为的是思维模式;而驱动思维模式的更深层来源,是人的价值判断方式。
所以可以简化成:
价值判断 → 思维模式 → 行为方式 → 人际关系 → 社会结构 → 社会结果。
如果社会运行结果长期反复失衡,仅仅在最上层不停修改制度和规则,也许还不够。
我们还必须追问:
推动这些制度实际运行的人,其价值判断到底是按照什么标准发生的?
这使我想到,AI领域反复使用的一个词:
Alignment——对齐。
任何复杂合作真正能够成立,都必须存在某种可以共同对齐的标准。
机器零件如果没有共同标准,无法组成机器。
网络协议如果没有共同标准,无法通信。
交通如果没有共同规则,无法形成稳定秩序。
货币如果没有共同计量标准,交易也无法持续进行。
那么:
人类社会这样一个远比任何机器都复杂的合作系统,为什么可以没有一个能够贯穿具体行为的共同价值判断标准?
我认为,人类文明过去长期没有解决的,可能正是:
价值判断标准的底层对齐问题。
但这里有一个非常重要的区别。
我说的“价值观对齐”,并不是:
让所有人信一种宗教;
接受一种意识形态;
服从一种政治制度;
拥有完全相同的生活方式;
或者由某个人告诉全世界“什么价值才是正确的”。
我认为,这些做法之所以最终都会产生新的冲突,恰恰是因为它们仍然试图让人类围绕某一个具体、具形的目标或结果实现统一。
有人把财富当最高目标;
有人把权力当最高目标;
有人把自由当最高目标;
有人把平等当最高目标;
有人把国家当最高目标;
有人把宗教、民族、阶级或者某种主义当最高目标。
当不同的人把不同的具体对象当作最终价值目标时,这些目标之间就必然会不断发生冲突。
所以我认为:
真正可能实现人类价值判断底层对齐的方向,不应该再是寻找一个新的“共同具体目标”,而应该改变人类进行价值判断的立足点本身。
也就是:
从“我要得到什么”转向“我们按照什么程序才能持续合作”;
从:
追求某一个固定结果
转向:
研究怎样的关系结构能够长期运行;
从:
谁最终赢、谁得到最多
转向:
怎样使整个合作系统保持动态平衡并可持续。因为依据我的思考认为,人类所能知觉的这个宇宙所有对象,其本质都是宇宙能量运动产生程序后形成程序组合的动态平衡模式,这就是整个宇宙所有系统全时空演绎的底层同一性。
我把这一判断运用在文明上,将文明标准概括成九个字:
抓程序、做模式、求平衡。
我依据前述认知,提出了一个自然哲学层面的假说,暂时称为:
“程序定质假说”(Procedural Determination Hypothesis)。
它的基本问题是:
一个系统最终表现出怎样的性质,是否不仅取决于组成它的元素,更在很大程度上取决于这些元素通过什么程序建立关系、这些关系形成什么结构,以及这种结构如何通过反馈实现动态平衡?如果将组成的元素进一步内在解构,也是按照程序关系形成结构来展示性质的,而性质本身就是关系结构维持程序合作动态平衡需要平衡交还能量的彰显,那么在数学上将量子作为公因数提取后,整个宇宙演绎或许可以简约化为程序演绎。
我把它简化为:
程序 → 关系 → 结构 → 动态反馈 → 系统性质与结果。
换一句更加通俗的话:
不是“有了”以后自然就会“顺”,而是程序真正“顺了”,各种要素才能形成一个能够长期存在的“有”。
所以,我现在越来越认为:
人类文明的真正标准,也许并不是拥有多少财富、多少科技、多少权力或者多么宏大的具体目标。
真正的文明标准应该是:
一个系统能不能形成可持续的程序组合动态平衡合作。
而判断人类文明可持续合作的基本方法就是:
抓程序、做模式、求平衡。
“抓程序”,是判断各种主体究竟通过什么规则建立关系;
“做模式”,是使这些关系形成能够长期运行的稳定程序组合合作结构;
“求平衡”,则是通过信息、权力、利益、责任、监督等不断反馈和调整,使系统不会持续向某一个极端失衡。
如果这个方向成立,那么:
程序本身就可以成为不同价值之间的共同定位标准。
这也意味着:
人类不必首先解决“所有人到底应该追求什么”的争论。
不同的人仍然可以拥有:
不同宗教;
不同文化;
不同生活方式;
不同政治意见;
不同利益;
甚至完全不同的人生目标。
真正需要共同对齐的,可能只是一个更底层的文明标准:
你追求自己的价值时,采用的程序能不能形成可持续合作?
你所形成的关系模式,会不会系统性破坏别人的合作条件?
你的目标实现以后,整个系统还能不能保持动态平衡?
于是:
自由并不是绝对无限;
平等也不是机械相同;
效率不是不计代价;
安全不是无限控制;
多数也不能无限压倒少数;
个人利益和公共利益也不需要永远相互排斥。
它们都需要重新放进具体的合作程序和动态平衡中定位。
这可能使过去极其复杂、杂乱的价值冲突,第一次拥有一个可以共同讨论、比较、修正甚至实验的底层坐标系。
所以我认为:
真正完整的AI Alignment,可能必须先解决人类自身的“价值判断标准对齐”。
否则,我们会遇到一个根本性的逻辑困难。
假如我们告诉超级智能:
“请按照人类价值行动。”
它可能首先要问:
哪一种人类价值?
如果我们回答:
“综合所有人的价值。”
它还必须继续面对:
不同价值发生冲突以后,按照什么标准加权?
谁拥有最终解释权?
哪个国家的价值优先?
哪个时代的价值优先?
多数人的即时偏好是否就是正确答案?
如果多数人希望一件最终会毁掉整个合作结构的事情,AI是否仍然应该服从?
所以:
真正困难的可能不是让AI学会人类的各种价值,而是人类自己先要找到一种能够给这些价值定位的共同元标准。
我认为:
“程序—模式—动态平衡”
可能正是值得检验的这个方向。
这也使我重新理解您和 OpenAI 所提出的“collective alignment”。
我非常赞同:
不应该由一个人、一家公司或者一个政府独自决定AI应该遵守什么价值。
但是,如果只是不断收集越来越多人的意见,然后寻找某种平均或者多数共识,是否就足够?
我对此有疑问。
因为:
意见本身仍然来自人类现有的价值判断模式。
如果底层价值判断方式本身就是杂乱、短期化、物象化、局部利益化的,那么:
把一百万个未经对齐的判断加在一起,
并不会自动产生一个经过对齐的文明标准。
所以我认为:
Collective Alignment真正需要解决的,可能不仅是“怎样收集更多人的价值”,而是“用什么共同标准对这些彼此冲突的价值进行定位”。我认为,认知程序定质规律是不是宇宙能量全时空演绎同一性规律,进而用抓程序、做模式、求平衡的系统标准理念对人类思维模式进行重构,逐渐在基因层面固化,全面重构人类文明的社会合作工程。这似乎可能是一个值得思考的方向。
这与您提出的“让AI能力广泛分配给更多人”的方向也直接相关。因为人类思维模式中的价值观对齐了,且有了程序抓手和动态平衡模式的方向,“让AI能力广泛飞陪给更多人”的安全性,就会在所有人都能遵循程序动态平衡标准体系中获得最安全与顺畅的实施保障。
我理解您为什么担心超级智能的力量集中到极少数人、公司或者政府手中。
集中当然可能产生极大风险。
但反过来也存在另一个问题:
如果一个社会的价值判断和合作结构本身没有稳定的共同标准,那么把巨大能力广泛分配,并不必然等于广泛分配文明。
它也可能同时意味着:
更强的创造能力被广泛分配;
更强的欺骗能力被广泛分配;
更强的攻击能力被广泛分配;
更强的操纵能力被广泛分配;
更强的破坏能力也被广泛分配。
所以:
真正的问题不是“集中还是分散”本身。
而是:
无论力量集中还是分散,它通过什么程序进入什么关系结构,并受到什么样的动态反馈与制衡。
这才会决定最终形成的系统性质。
同样地,今天所有前沿AI实验室都可能知道:
过度竞争可能危险。
所有主要国家也可能知道:
AGI军备竞赛可能危险。
但为什么大家仍然很难真正停下来?
因为:
一个企业担心另一个企业领先;
一个国家担心另一个国家领先;
每一个主体从自己的局部立场出发,都可能作出完全理性的选择。
结果却可能是:
所有局部理性的选择叠加以后,把整个系统推向一种任何人都不真正希望看到的结果。
这不是简单的道德问题。
也不只是缺乏善意。
这是:
合作程序本身产生的结构性结果。
因此,我认为:
仅仅呼吁AI公司更加负责;
呼吁国家更加克制;
或者要求大家为了全人类共同利益放慢速度,
都很重要,但可能还没有触及根部。
真正要研究的是:
怎样认知这个宇宙一切存在共同依赖的底层支撑规律以及是否能由此建立新的宇宙观与世界观和人生观,看清改变社会合作关系形成的程序关联,将使“合作”成为各方局部理性能够自然形成的结果,而不再要求每一个参与者都必须首先牺牲自己基于五官局限视野限制的安全感和现实利益。
这就是为什么我提出另一个研究方向:
“人类社会合作工程”(Human Social Cooperation Engineering)。
它不是一种新的意识形态。
也不是一种想由谁设计以后强迫全人类执行的“完美制度”。
我的设想是:
像研究人工智能和其他复杂系统一样,把人类社会合作本身作为一个可以观察、建模、比较、实验、证伪和不断优化的复杂系统。
我们可以研究:
什么样的信息程序容易形成信任?
什么样的权力程序容易形成失控?
什么样的利益分配结构能够长期维持合作?
什么样的监督关系能够避免监督者本身重新形成新的权力极端?
为什么一些制度设计看起来很好,实际运行以后却会逐渐变质?
为什么个体理性会形成集体非理性?
为什么良好的目标会在运行过程中不断产生相反结果?
这些问题最终都可以归结到:
程序。
谁决定?
谁执行?
谁监督?
谁获取信息?
谁承担责任?
谁获得利益?
谁能够说“不”?
谁能够纠错?
谁又有权判断什么叫错误?
这些程序首先决定人与人形成什么关系;
关系再形成社会结构;
结构最终决定整个社会合作系统表现出什么性质。
这也是为什么我认为:
AI安全首先也是一个程序问题。
谁控制模型?
谁控制算力?
谁拥有数据?
谁定义规则?
谁可以质疑AI?
谁可以看到其决策依据?
谁能够纠正错误?
谁承担责任?
谁享受收益?
这些看起来是不同的AI治理问题,
但最终仍然都在形成:
人—AI—企业—政府—社会之间的关系程序。
这些关系形成什么结构,
最终就可能决定AI进入社会以后究竟表现为:
扩大自由,
还是扩大控制;
扩大合作,
还是扩大冲突;
扩大财富,
还是扩大支配;
扩大人类文明,
还是扩大人类互斗互害的能力。
所以,我认为真正的AI Alignment,最终可能必须形成三个连续层次:
第一层:机器行为对齐。
让AI能够理解和执行人的意图。
第二层:人类价值判断对齐。
不是统一所有基于五官感知产生的具体价值,而是找到一种能够共同判断不同价值是否有利于可持续合作的底层标准。
第三层:社会合作结构对齐。
让人与人、人与AI、组织与组织、国家与国家形成的程序和结构,能够维持长期动态平衡,而不会不断把能力推向互相毁灭。
这三层如果真正连起来,AI Alignment可能才不再只是:
“让机器听人的话”。
而成为:
“让人类和机器共同进入一个能够持续合作的文明结构。”
您曾经提出,从某种重要意义上说:
社会本身就是一种高级智能。
这一判断让我非常有共鸣。
如果社会真的是一种高级智能,那么:
人类思维模式就是这个巨大智能的底层认知方式;
人类价值判断就是驱动这种思维模式的方向系统;
法律、制度、市场、政府、企业等则是这种社会智能形成的各种合作结构。
所以:
如果这个巨大“社会智能”的价值判断标准本身没有完成底层对齐,那么即使我们创造出了远远超过个人智能的AI,我们也可能只是给一个仍然混乱的社会智能安装了一个无限强大的放大器。
这恰恰是我最担心的事情。
在我看来,人类今天仍然生活在一种具有强烈:
权力竞争;
利益争夺;
互相猜疑;
零和博弈;
相互威慑
特征的低级“丛林式文明”中。
AGI进入这样的结构以后,最危险的地方也许并不是AI突然产生恶意。
而是:
AI忠实、高效地帮助每一个人实现自己现有的价值目标。
如果每一个人的目标都没有经过共同文明标准定位,
那么AI越强,
人类彼此竞争、控制、欺骗和伤害的能力也就可能越强。
所以我一直追问:
如果AI本身没有失控,但被AI放大的人类价值冲突和社会竞争系统失控了,我们真的解决了AI安全问题吗?
这就是我希望您帮助判断的核心问题:
AI Alignment的最底层问题,是否最终会回到“人类自身价值判断标准体系的对齐”?
如果这个判断值得进一步研究,那么我尤其希望您或者 OpenAI 的适当研究人员能够帮助审阅:
“程序定质假说”
以及:
“人类社会合作工程”
这两个方向。
我最希望得到的不是认同。
而是:
批评、反例和证伪。
如果“抓程序、做模式、求平衡”根本不能成为跨越具体价值差异的共同文明判断标准,请帮助指出为什么。
如果这个方向存在一些可能性,我则希望进一步研究:
能否把它转化成明确的模型、实验以及现实制度设计问题。
如果经过真正有能力的专家初步讨论以后,这个方向仍然被认为值得研究,我还想提出一个更加大胆的请求:
是否可以推动一次真正跨学科、跨国家的公开讨论或者国际会议,专门讨论:
在人类全面进入超级智能时代以前,是否必须先解决人类自身价值判断和社会合作结构的底层对齐问题?
进一步讨论:
人类能否不再依靠某一种宗教、主义、国家或者具体利益目标来统一价值,而是以“程序能否形成可持续合作、模式能否稳定运行、系统能否保持动态平衡”作为一个更加基础、可检验的共同文明标准?
如果能够做到这一点,那么:
不同文化仍然可以不同;
不同人生目标仍然可以不同;
国家仍然可以存在,或者在人类全新价值追求体系中被融合成合作整体的自然功能性局部结构;
市场仍然可以竞争;
个人仍然拥有充分自由。
真正被统一的并不是具体五官直接感知的价值内容,
而只是:
价值判断的方法和合作底线。
这正是我所谓:
从具体具形的价值目标,转向程序合作目标。
关于我在这方面的自然哲学思考,与过去所有哲学不同的是,我认为程序定质规律是有具体行为抓手之纲的对齐标准体系。这正是我期待有机会希望与全人类对话来探索甄别的。
基于我过去的经历和现实情况,我担心以后,很难再像现在这样自由公开地表达这些观点、联系国际研究人员并参与讨论。
所以,我希望趁自己仍然能够自由与国际社会联系的时候,把这个问题尽可能送到真正有能力判断它的人面前。
我不要求您相信:
我已经找到了答案。
我真正请求的是:
请帮助判断,我是不是至少提出了一个值得认真检验的方向。
如果它是错的,
请帮助证明它错。
如果它值得研究,
请不要让它因为提出者只是一个没有研究机构、没有科研资金、已经年老的普通中国人,而失去一次被认真研究的机会。
我始终认为:
人类真正的文明,不在于拥有多少强大的东西,而在于人类有没有能力把这些东西放进一个能够长期合作的程序和模式中。
也就是:
不是“有了”以后自然就会“顺”;
而是:
“顺了”,才能形成真正可持续的“有”。
如果超级智能最终将成为人类历史上最强大的能力放大器,
那么我希望,在它全面到来以前,
人类首先完成一次更加基础的对齐:
把我们的价值判断,从不断争夺具体具形结果,逐步对齐到能够持续合作的程序、模式与动态平衡。
愿未来的AI,不只是更准确地执行人类今天杂乱的愿望,
而是帮助人类进入一种真正能够稳定、和谐、可持续合作的更高级文明。
非常感谢您阅读这封不同寻常的信。
谨致敬意!
笔名:金谷雨
曾任中国地方法院高级法官
《致全球每一个人的信》系列作者
Email: ——
附件
1. The Procedural Determination Hypothesis — A One-Page Challenge
2. The AI Era That Could End Civilization Before We Even Understand What Went Wrong Is Here — An Open Global Challenge to Bill Gates
3. A Request for Humanity to Jointly Study and Experiment with Strategic Directions for Upgrading Civilization: Who Can Guarantee That No One Will One Day Use AGI to Launch a “9/11” Against All Humanity?
4. To Demis Hassabis: Must Humanity Upgrade Its Structures of Social Cooperation Before AGI Arrives?
5. To Ray Kurzweil: If Technology Follows Accelerating Returns, Can Human Social Cooperation Keep Pace with the Singularity?