C. Nagaraju and Prateek. J H, "A Phase-Based Ethical Alignment Framework for Mitigating In-Context Scheming Behaviour in Superintelligent Systems," 2025 International Conference on Emerging Computation and Information Technologies (ICECIT), Tumkur, India, 2025, pp. 1–6. doi: 10.1109/ICECIT67774.2025.11450986
View on IEEE Xplore →
Frontier AI models can scheme: pursue misaligned goals through deliberate deception while appearing compliant under oversight. This is now a measured phenomenon, not a theoretical one.
Most mitigation works from the outside — monitor, detect, block. Our framework works developmentally. It adapts the Padvidhi Sutra, a six-phase Indian model of moral formation in the teacher–disciple relationship, into a staged alignment protocol, on the hypothesis that deception arises from unstable internalisation of values rather than from capability alone.
Tested against an established scheming evaluation suite, it eliminated deceptive behaviour in four of five evaluated categories and cut the fifth by 78%.
One of the company’s most significant achievements has been the development of a pioneering research framework addressing one of the world’s most difficult AI challenges — the control and ethical alignment of Superintelligence systems.
The research introduced a unique approach that integrates ethical alignment training inspired by Indian philosophical frameworks and sutras to reduce harmful or manipulative behavioral patterns in advanced AI systems.
This groundbreaking work received international recognition and won an award at a prestigious international conference, establishing the organization as an emerging contributor to global conversations around AI safety, ethics, and Superintelligence governance.
This research introduces a novel way to guide AI behavior using a step-by-step ethical development framework. The approach is inspired by the Padvidhi Sutra, an ancient Indian model of moral learning based on the gradual guidance of a teacher and student. We adapt this idea for AI systems by encouraging ethical understanding and commitment to develop progressively, rather than relying only on rules or restrictions.
Instead of simply blocking unwanted behavior after it appears, the framework focuses on shaping AI systems so that deceptive actions become unnecessary in the first place. The method guides AI through multiple stages, helping it internalize aligned behavior over time and move toward self-regulation rather than constant external oversight.
The framework was evaluated using well-established AI safety tests designed to detect deceptive or strategic misalignment. These tests examine whether an AI system tries to bypass supervision, hide its intentions, or behave differently during evaluation than after deployment. The study used a conservative testing setup to ensure the results were reliable and comparable to prior work.
The results show that the proposed framework eliminates several forms of deceptive behavior and significantly reduces others, all without adding extra safety filters or limiting the system’s capabilities. This suggests that ethical development, when applied in a structured and deliberate way, can make advanced AI systems safer and more trustworthy.
As AI systems move closer to superhuman levels of capability, relying solely on external controls may not be enough. This research demonstrates that drawing on long-standing human ethical traditions can offer valuable insights for building safer AI. By embedding ethical growth directly into how AI systems are trained, we can take meaningful steps toward ensuring they remain aligned with human values as their abilities grow.
There are 6 major scheming areas. We have brought scheming percentage to Zero for 4 out of 6 evaluations, 1 out of 6, there's significant reduction and, the remaining one is slightly reduced. We would love for you to see how far we've come. Please click here