Chinese Edition: 关于人工智能对齐的数学思考
What makes neural networks powerful also makes alignment difficult:
Inverse learning: learns from data, but produces jagged intelligence.
Propertyless representations: learn across any modality, but are also “lawless.”
Full-rank transformations: connect relationships across any data, but do not provide moral or or principled separation.
What Shakespeare would say about neural networks and AI alignment ?
What Must Be Aligned?
At a high level, alignment concerns what the system is trying to achieve, how it interprets human instructions, how it uses its learned capabilities, and how it behaves in the world. Its goals, outputs, decisions, and actions are expected to remain consistent with physical reality, legal requirements, safety constraints, institutional policies, social norms, and human values. These layers are not identical. Physical laws constrain what is possible, laws and policies constrain what is permitted, and social norms and human intentions constrain what is considered appropriate.
Alignment for an Ill-Posed Inverse Problem
“Inverse problems are typically ill-posed, as opposed to the well-posed problems usually met in mathematical modeling. Of the three conditions for a well-posed problem suggested by Jacques Hadamard (existence, uniqueness, and stability of the solution or solutions) the condition of stability is most often violated.” (Mathematical and computational aspects of “Inverse Problem”
AI systems are not given a complete description of acceptable behavior. They infer it from examples, instructions, corrections, rankings, policies, and observed consequences. This inverse-learning ability is a major source of AI’s power because it allows the system to discover structures that were never fully written down. But the same process makes alignment ill-posed. Different learned structures may fit the same feedback, human requirements may conflict, and small changes in context may produce different behavior. We cannot directly control the internal solution; what we can control are the boundary conditions under which learning and inference take place.
Propertyless and “Lawless” Neural Computation
Neural networks operate through numbers, vectors, activations, and transformations. These elements are not intrinsically truthful, safe, fair, harmful, or deceptive. Nor do they inherently obey physical, legal, social, or moral laws. This propertyless (Deep Manifold Part 2: Neural Network Mathematics, Section 2.2.4 ) and lawless character gives neural networks extraordinary flexibility across domains. But it also means that human values are not built into individual neurons, parameters, or features. They can influence behavior only through learned relationships and imposed boundary conditions.
Full-Rank Relational Structure
From a category-theoretic perspective, neural networks learn relationships, transformations, and compositions rather than isolated objects. Full-rank transformations preserve many independent relational directions, allowing knowledge and capabilities to remain densely connected and recompilable. This connectedness is another source of AI’s power. But helpful and harmful behaviors may share the same representations, pathways, and relationships. Full rank preserves relational freedom; it does not provide moral separation.
Jagged Neural Fixed-Point Field Makes Alignment Harder
AI alignment is made more difficult by the geometry of today’s neural fixed-point field. In a rolling, piecewise-smooth field, inference can traverse relatively coherent pathways. Prompts, reasoning steps, tool use, memory access, observation, and final response can remain connected through a more stable iterated-integral path. Such a field does not guarantee alignment, but it makes aligned behavior easier to sustain because the computational pathway is more continuous, less fragile, and less dependent on repeated correction.
By contrast, a jagged and highly corrugated fixed-point field makes traversal unstable. Steep peaks and deep valleys increase drift, backtracking, retry loops, replanning, and diverging routes. In this setting, reasoning may wander, tool use may become inconsistent, memory retrieval may become less reliable, and the final behavioral outcome may become more sensitive to small perturbations. Alignment is therefore harder not simply because values are difficult to define, but because the computational field through which behavior is formed is itself unstable.
This matters especially for agentic workload. Alignment is not only a matter of a single output token or isolated answer. It must persist across a sequence of actions, observations, reasoning steps, and tool interactions. When the fixed-point field is jagged, the agent’s path becomes more fragile, slower, and more corrective. The system may remain locally capable, yet globally less trustworthy. In this sense, jagged geometry increases alignment difficulty by making inference traversal itself less stable, less coherent, and less controllable.
The Mathematical Foundation for AI Alignment
Deep Manifold provides a mathematical foundation for understanding what AI alignment is actually trying to influence. It distinguishes mathematical covers, formed by weights, from physical covers, formed by activations. The mathematical covers organize the network’s relational structure, while the physical covers make that structure observable during computation. Together, their successive organization forms the neural fixed-point field in which learned behavior becomes possible.
Within this framework, inference is understood as an iterated-integral process through successive covers. Each layer transforms and accumulates the relational structure produced by earlier layers, while the input, prompt, policies, feedback, and context act as boundary conditions on the resulting pathway. Alignment therefore does not directly control one neuron, feature, circuit, or hidden location. It attempts to influence the boundary conditions, covers, fixed-point field, and global inference pathways through which behavior is formed.
This connects AI alignment to the underlying mathematics of neural-network learning and computation. Before proposed remedies can be evaluated, we must first understand the mathematical system within which aligned and misaligned behavior emerge.
Alignment Across the Entire AI Lifecycle
Beyond external mechanistic interpretability and alignment applied after the fact, often near the end of training, AI alignment should be considered from every direction and across the entire model lifecycle.
Architecture already determines the admissible learning space, relational structure, and computational pathways, as discussed in Mathematical Considerations for Neural Network Architecture.
Training progression then determines how those structures are formed, deformed, and stabilized over time, as discussed in Mathematical Considerations for Training Progression.
Alignment should therefore not be treated as a final corrective layer. Its mathematical elements must be considered in architecture, data, boundary conditions, training progression, inference, evaluation, and deployment.
This article is now listed under Mechanistic Interpretability & Alignment and Mathematical Considerations Series in Deep Manifold, Two Years Later: 2024–2026.









