The Third Flag · Accountable Autonomy

Intelligence is not a parameter count.
It is an architecture.

A third position for agentic AI — not contain it, not command it. An agent may act. That does not make every action legitimate. Accountable autonomy is the principle; the ORDIS Agentic Plane is the working experiment.

V2 COMPLETE
10/10 Graduation Gate
0 False Trusted Outcome
LOCAL core, independent
I — THE FALSE BINARY

Two flags, one false choice

The public debate about AI keeps collapsing into two positions.

AI is a threat contain it. / AI is a tool command it.
// both treat "what happens once an agent can act" as already settled

Neither position engages the actual engineering question: the moment a system can execute code, call tools, and persist memory, it has agency. Containment denies that agency is already here. Command assumes agency is automatically safe as long as a human nominally holds the leash. We think there is a third way — not a compromise between the two flags, but a different question entirely: given that an agent can act, what makes any one of its actions trustworthy?

II — AGENCY IS NOT AUTHORITY

The ability to act is not the right to act

Agency Authority

An AI can move. It can write a file, call an API, spend a budget, send a message. None of that implies the action was admitted — that someone or something with the standing to decide said yes, under a scope it can be held to. In our system, every action is admitted under an execution class it cannot choose for itself, and a class can never be silently downgraded to something cheaper or less scrutinized than what the task structurally requires. Agency is the engine. Authority is the steering. An engine without steering is not autonomy — it is just motion.

III — EXECUTION IS NOT TRUTH

Running is not proof of being right

Execution Verification

A task that completed and returned "done" has told you one thing: it stopped running. It has not told you it was correct. Most systems treat completion as success. We treat completion as a claim — one that must be backed by evidence a separate verifier actually checks before it is allowed to count as true. A passing test receipt moves a claim toward verified. A self-report does not. This is the difference between a system that finishes tasks and a system whose claim of finishing can be trusted.

IV — TRUST MUST BE EARNED

Evidence, then proof, then verification

Evidence Proof Verification

Trust is not a starting condition; it is the output of a pipeline. Every run accumulates evidence as it executes. That evidence is assembled into a proof bundle. The proof bundle is checked — for completeness, for contradiction, for whether it actually supports the claim being made — before a run is allowed to be called verified. A run with no evidence, weak evidence, or contradicted evidence stays unverified, no matter how confidently it reports its own success. Confidence is not evidence. It never was.

V — MEMORY IS GOVERNANCE

What a model says is not what a system remembers

Memory Model Output

A model can say almost anything with the same fluent confidence, whether it is right or not. If that output can walk straight into long-term memory, the system's memory becomes exactly as reliable as its most confident hallucination. We treat memory promotion as a governed act, not a side effect of generation: a claim becomes canonical memory only after it is backed by a fully verified proof, bound digest-for-digest to the memory record itself. Memory that no longer matches its proof is stale, contradicted, or revoked — never quietly kept because deleting it would be inconvenient.

VI — FAILURE MUST BE LEGIBLE

An interruption is a question, not a retry button

Recovery Replay

Systems fail. Processes restart. Networks drop mid-request. The dangerous failure mode is not the crash — it is what happens next. A side effect that may already have happened (a file written, a message sent, a payment made) must never be blindly repeated on the assumption that "retry" is always safe. An ambiguous side effect is escalated for a real decision. A clean interruption with no observed side effect is a genuine retry candidate. The difference between those two cases is not cosmetic — collapsing it is how systems double-charge, double-post, and double-act. We reconcile. We do not replay blind.

VII — AUTONOMY WITH ACCOUNTABILITY

The principle, made operational

This is the synthesis of I–VI, and where the manifesto stops being philosophy and starts being a testable claim.

Not absolute sovereignty. The local core is independently operational without any network dependency. An external expert model is optional and policy-gated, consulted only for the heaviest, most explicitly-scoped tasks — never a silent requirement the interactive loop secretly depends on. We say "independent," not "total," because only one of those words is actually true.
Not a claim of universal safety. We do not say the system cannot be broken. We say what we tested, how we tried to break it, and what held: a ten-criterion Graduation Gate attacking identity, admission, execution, evidence, proof, recovery, memory, skill, workspace, and gateway together — plus a creative red-team pass beyond those ten. Both are public. Both report their real result, not a marketing summary of it.
Evidence over narration. Every claim in this manifesto has a corresponding, checkable artifact — a test, a ledger entry, a tagged release — in the system it describes. If a claim here ever stops matching what the code actually does, the claim is wrong and gets corrected, not the code silently reinterpreted to fit the prose.

The Third Flag is the thesis.

Accountable autonomy is the principle. The ORDIS Agentic Plane is the experiment. Evidence decides whether the experiment succeeds.

— ORDIS
Engineered by Uğur Can Leba · Agentic Plane, V2 COMPLETE.
Üçüncü Bayrak · Hesap Verebilir Özerklik

Zekâ bir parametre sayısı değildir.
Bir mimaridir.

Ajan yapay zekâ için üçüncü bir konum — ne kısıtla, ne komuta et. Bir ajan hareket edebilir. Bu, her hareketinin meşru olduğu anlamına gelmez. Hesap verebilir özerklik ilkedir; ORDIS Agentic Plane çalışan deneydir.

V2 TAMAMLANDI
10/10 Graduation Gate
0 False Trusted Outcome
LOCAL çekirdek, bağımsız
I — YANLIŞ İKİLEM

İki bayrak, tek yanlış seçim

Yapay zekâ üzerine kamusal tartışma sürekli iki konuma sıkışıyor.

Yapay zekâ bir tehdit kısıtla. / Yapay zekâ bir araç komuta et.
// ikisi de "bir ajan hareket edebildiğinde ne olur" sorusunu çoktan çözülmüş sayar

Hiçbiri asıl mühendislik sorusuyla ilgilenmiyor: bir sistem kod çalıştırabildiği, araç çağırabildiği ve bellek kalıcılaştırabildiği an, onun agency'si (eyleyicilik) vardır. Kısıtlama, bu agency'nin çoktan burada olduğunu reddeder. Komuta, bir insan sözde ipi elinde tuttuğu sürece agency'nin otomatik olarak güvenli olduğunu varsayar. Bizce üçüncü bir yol var — iki bayrak arasında bir uzlaşma değil, tamamen farklı bir soru: bir ajan hareket edebiliyorsa, onun hareketlerinden hangisini güvenilir kılan nedir?

II — AGENCY AUTHORITY DEĞİLDİR

Hareket edebilme, hareket etme hakkı değildir

Agency Authority

Bir yapay zekâ hareket edebilir. Dosya yazabilir, API çağırabilir, bütçe harcayabilir, mesaj gönderebilir. Hiçbiri, hareketin kabul edildiği anlamına gelmez — kararı verecek konumdaki birinin ya da bir şeyin, hesap verebileceği bir kapsam altında evet dediği anlamına gelmez. Bizim sistemimizde her hareket, kendisinin seçemeyeceği bir execution class altında kabul edilir ve bir sınıf, görevin yapısal olarak gerektirdiğinden daha ucuz ya da daha az denetlenen bir şeye asla sessizce düşürülemez. Agency motordur. Authority direksiyondur. Direksiyonsuz bir motor özerklik değildir — sadece harekettir.

III — EXECUTION TRUTH DEĞİLDİR

Çalışmış olmak, doğru olduğunun kanıtı değildir

Execution Verification

Tamamlanmış ve "bitti" dönmüş bir görev size tek bir şey söylemiştir: çalışmayı durdurdu. Doğru olduğunu söylememiştir. Çoğu sistem tamamlanmayı başarı sayar. Biz tamamlanmayı bir iddia sayarız — doğru sayılmasına izin verilmeden önce ayrı bir doğrulayıcının gerçekten kontrol ettiği kanıtla desteklenmesi gereken bir iddia. Geçen bir test makbuzu bir iddiayı doğrulanmışa yaklaştırır. Kendi kendine rapor etmek yaklaştırmaz. Bu, görevleri bitiren bir sistemle, "bitirdim" iddiasına güvenilebilecek bir sistem arasındaki farktır.

IV — GÜVEN KAZANILMALIDIR

Önce kanıt, sonra proof, sonra doğrulama

Kanıt Proof Verification

Güven bir başlangıç koşulu değildir; bir hattın çıktısıdır. Her run, çalışırken kanıt biriktirir. O kanıt bir proof bundle'a derlenir. Proof bundle — bütünlük, çelişki ve iddia ettiği şeyi gerçekten destekleyip desteklemediği açısından — bir run'ın doğrulanmış sayılmasına izin verilmeden önce kontrol edilir. Kanıtı olmayan, zayıf ya da çelişen kanıtı olan bir run, kendi başarısını ne kadar güvenle rapor ederse etsin doğrulanmamış kalır. Güven kanıt değildir. Hiçbir zaman olmadı.

V — BELLEK YÖNETİŞİMDİR

Bir modelin söylediği, bir sistemin hatırladığı değildir

Memory Model Output

Bir model, doğru olsun olmasın, hemen hemen her şeyi aynı akıcı güvenle söyleyebilir. O çıktı doğrudan uzun-vadeli belleğe girebiliyorsa, sistemin belleği tam olarak en özgüvenli halüsinasyonu kadar güvenilir olur. Biz memory terfisini üretimin bir yan etkisi değil, yönetişimli bir eylem olarak ele alırız: bir iddia, ancak tam doğrulanmış bir proof'la desteklendiğinde ve memory kaydına digest-digest'e bağlandığında canonical memory olur. Proof'uyla artık eşleşmeyen bellek eski, çelişkili ya da geri alınmıştır — silmek elverişsiz olduğu için sessizce tutulmaz.

VI — HATA OKUNABİLİR OLMALI

Bir kesinti bir sorudur, bir yeniden-dene düğmesi değil

Recovery Replay

Sistemler bozulur. Süreçler yeniden başlar. Ağlar isteğin ortasında düşer. Tehlikeli olan başarısızlık kipi çökme değildir — sonra ne olduğudur. Zaten olmuş olabilecek bir yan-etki (yazılmış bir dosya, gönderilmiş bir mesaj, yapılmış bir ödeme), "yeniden dene her zaman güvenlidir" varsayımıyla asla körü körüne tekrarlanmamalıdır. Belirsiz bir yan-etki gerçek bir karar için yükseltilir. Gözlenen bir yan-etkisi olmayan temiz bir kesinti gerçek bir retry adayıdır. Bu iki durum arasındaki fark kozmetik değildir — onu birleştirmek, sistemlerin çift-ücretlendirdiği, çift-gönderdiği ve çift-hareket ettiği yerdir. Biz uzlaştırırız (reconcile). Körü körüne tekrarlamayız.

VII — HESAP VEREBİLİR ÖZERKLİK

İlke, operasyonel hâle gelmiş

Bu, I–VI'nın sentezidir; manifestonun felsefe olmaktan çıkıp test edilebilir bir iddiaya dönüştüğü yer.

Mutlak egemenlik değil. Yerel çekirdek, hiçbir ağ bağımlılığı olmadan bağımsız çalışır. Harici bir uzman model isteğe bağlı ve politika-kapılıdır, yalnızca en ağır, açıkça sınırlandırılmış görevler için danışılır — etkileşim döngüsünün gizlice bağımlı olduğu sessiz bir gereklilik asla değil. "Bağımsız" diyoruz, "total" değil — çünkü bu iki kelimeden yalnızca biri gerçekten doğru.
Evrensel güvenlik iddiası değil. Sistemin kırılamayacağını söylemiyoruz. Neyi test ettiğimizi, onu nasıl kırmaya çalıştığımızı ve neyin tutunduğunu söylüyoruz: kimlik, kabul, execution, kanıt, proof, recovery, memory, skill, workspace ve gateway'e birlikte saldıran on kriterlik bir Graduation Gate — artı bu on kriterin ötesinde yaratıcı bir red-team geçişi. İkisi de kamuya açık. İkisi de pazarlama özeti değil gerçek sonucunu raporluyor.
Anlatı değil kanıt. Bu manifestodaki her iddianın, anlattığı sistemde karşılık gelen, kontrol edilebilir bir artefaktı vardır — bir test, bir defter girdisi, etiketlenmiş bir sürüm. Buradaki bir iddia, kodun gerçekte yaptığıyla eşleşmemeye başlarsa, iddia yanlıştır ve düzeltilir — kod sessizce metne uyacak şekilde yeniden yorumlanmaz.

Üçüncü Bayrak tezdir.

Hesap verebilir özerklik ilkedir. ORDIS Agentic Plane deneydir. Deneyin başarılı olup olmadığına kanıt karar verir.

— ORDIS
Uğur Can Leba tarafından inşa edildi · Agentic Plane, V2 TAMAMLANDI.