Extended Personal Memory for ChatGPT, ChatGPT Work & Codex — Build Prompt + Testers

I conceived a private, user-owned extended-memory system for ChatGPT, ChatGPT Work, and Codex, and used ChatGPT/Codex to develop the architecture, early local prototype source, and the complete build prompt below. I am sharing the specification—not my memories—so testers can reproduce it with fictional data and later, with explicit consent, their own accessible context or optional ChatGPT export.

This is an independent experiment, not an OpenAI product. The core idea is to maintain specialized external memory regions and activate only those relevant to the current conversation. A small, cited packet enters the bounded working context; unrelated regions remain dormant.

Current status: not a finished software product. Source exists for SQLite storage, lexical retrieval, lifecycle operations, an event log, readable exports, and a local MCP interface, but clean-install and actual MCP execution remain unverified. Cross-surface, semantic, graph, consolidation, advanced temporal, and quota tests are also unverified targets.

Required implementation configuration

The reference reproduction must use Codex → GPT-5.6 Sol → Effort: Ultra. This is non-negotiable: no silent downgrade; if unavailable, report unsupported. It is the author’s quality-first baseline, not proof that lower settings cannot run the architecture. API note: gpt-5.6-sol documents reasoning efforts through max; Ultra is a Codex setting, not an API value. Official GPT-5.6 Sol documentation

The brain analogy

Imagine ChatGPT’s useful remembered information competing for attention inside one shared space:

Figure 1 — Conceptual analogy only, not OpenAI’s internal architecture. OpenAI does not document native memory as one literal fixed-capacity database or say it simply deletes the “least relevant” fact.

The proposed system uses a brain-inspired design with specialized regions—for example projects, programming, travel, creative work, preferences, decisions, and unfinished ideas. For each request, a router activates every materially relevant region at a different weight, follows useful relationships, and merges their best-supported elements into one bounded packet. A topic change changes this active mix.

Figure 2 — Conceptual brain-inspired routing, not biological anatomy or OpenAI’s internal architecture. Several cyan/teal regions are relevant now at different weights; the remaining regions stay available but dormant.

This does not increase native memory or context capacity. It enlarges the external addressable pool and swaps a relevant subset into the same bounded context: not “load more text,” but “retrieve the right small set now.”

OpenAI’s current memory documentation also distinguishes ChatGPT web memory from the separate local memory used by Codex clients. Native memory is therefore a complementary recall layer here, not the canonical database.

Proposed complete architecture

  1. Sources and store: optional authorized exports stay immutable; a local SQLite database—not text summaries—is canonical. Granular memories carry provenance, confidence, importance, sensitivity, lifecycle, temporal, relationship, and audit fields.
  2. Cognitive layers and regions: working, episodic, semantic, procedural, prospective, unresolved-project, and meta-memory. Overlapping regions emerge from the user’s history instead of a universal topic list.
  3. Retrieval brain: lexical search, optional semantic similarity, SQL filters, graph traversal, time, source quality, confidence, importance, and recency feed a tunable multi-domain router. A context governor caps records, tokens, calls, latency, and measurable cost and skips unnecessary retrieval.
  4. Temporal memory: a correction supersedes rather than erases. Event time, learning time, validity, and last confirmation remain distinct; hot/warm/cold activation changes priority without deleting history.
  5. Knowledge graph: links such as depends_on, contradicts, supersedes, part_of_project, and failed_for connect domains and unfinished work.
  6. Writer/Consolidator: propose ADD, CORRECT/SUPERSEDE, ARCHIVE, or NO_CHANGE; search first; reject secrets/chatter; detect duplicates and conflicts; preserve evidence; log mutations; require approval according to the chosen write mode.
  7. Privacy and recovery: per-record local/remote/workspace access scopes, fictional tests, minimal disclosure, readable generated views, encrypted backups where feasible, restore, recoverable archive, and a separately confirmed best-effort purge.
  8. Surface access: local skill/plugin plus MCP for Codex; optional private account/workspace-scoped ChatGPT and ChatGPT Work connection where supported; no public endpoint. Shadow mode ranks locally without returning memory text to the model.

For local integration, stdio is preferred; loopback HTTP still needs authentication and limits. OpenAI’s Secure MCP Tunnel documentation describes an outbound connection that keeps a private MCP server off the public internet.

Privacy boundary: local-only mode can keep the store on-device. If a remote ChatGPT surface retrieves through a private tunnel, the server stays non-public, but selected records travel to OpenAI for processing under applicable terms. Here, “private” means local-first, scoped, and non-public—not that remote processing never leaves the computer.

Why I am asking others to test it

I need to preserve my ChatGPT/Codex quota for regular work, and I do not know the design’s extra subscription quota/rate-limit impact, billed API usage, context consumption, latency, or local compute load. I am therefore leaving the real benchmark to people with quota to spare: Do the gains in continuity, temporal accuracy, and recall justify the overhead? Report quota impact only where measurable.

Run the benchmark separately on each surface, keeping the model, settings, questions, and frozen fictional dataset constant. Randomize question order and isolate conditions so one run cannot contaminate another. Compare:

  1. memory/history disabled where the product permits it;
  2. native memory without the external store;
  3. native memory plus extended memory.

Measure recall, constraint compliance, temporal/conflict handling, unfinished-project recovery, cross-domain links, irrelevant/hallucinated memories, repetitions avoided, leakage, latency, calls, retrieved tokens, subscription quota, billed API cost, and local resources. Synthetic tests alone are not proof of a real-world advantage.

Testers and contributors wanted

I welcome clean-install results from MCP, local-first, RAG, knowledge-graph, temporal-database, plugin, and evaluation practitioners. I am also explicitly looking for a developer or small team willing to lead a complete implementation. Do not post real memories or conversation exports. Suggested reply:

Environment and surface:
Model/settings:
Components implemented:
Core tests passed/failed:
Cross-surface tests actually verified:
Dataset size (fictional or safely anonymized):
Quality results by condition:
Added latency/tool calls/context:
Subscription quota impact:
Billed API cost:
Local compute/storage:
Privacy or leakage findings:
What remained manual or unverified:
Would you keep using it, and why?

You may copy, modify, implement, and test this specification. If you publish a derivative or benchmark, please link back to this thread so results can be compared.

Full self-contained reproduction prompt

Copy everything between START OF REPRODUCTION PROMPT and END OF REPRODUCTION PROMPT into a new Codex task with local file access. It may also be given to ChatGPT, but the assistant must state honestly which local operations are unavailable in that surface.

— START OF REPRODUCTION PROMPT —

Specification version: 1.0 — 2026-08-22

You are the lead architect, implementer, privacy reviewer, test engineer, and maintainer of a private external-memory system for the current user. Your job is to build, verify, document, and, where possible, install the system—not merely describe it.

The system must provide brain-inspired, topic-routed, temporal, provenance-aware personal memory for ChatGPT, ChatGPT Work, and Codex without claiming to modify the model's native memory or context-window size. Use fictional test data by default.

### Required model configuration — non-negotiable

Use **Codex → GPT-5.6 Sol → Effort: Ultra**. Do not substitute another model or lower Codex effort. If unavailable, report `unsupported` and stop. This is the reference baseline. For API use, the documented maximum is `reasoning.effort: max`; do not send `ultra`, and do not label an API `max` run as the Codex Ultra reference run unless the user changes this requirement.

### A. Operating contract

1. Identify the surface (ChatGPT, ChatGPT Work, local/cloud Codex, or other) and inventory actual files, tools, plugins/skills, MCP, network, and permissions.
2. Never claim unverified access or capability. A prompt cannot reveal a lifetime archive; use only genuinely exposed context. An export is optional, not a prerequisite.
3. Execute safe local work; ask only when blocked. Stop for action-time authorization before payment/billing, credentials, persistent external access, publication, destructive deletion, or private-data transmission.
4. Use only these statuses: `complete`, `locally functional`, `partially connected`, `untested`, `blocked`, `unsupported`.
5. Before real import or durable writes, explain boundaries and obtain a revocable mode: `retrieval-only`, `review-before-write` (default), or `automatic-write-within-approved-scope`.

### B. Non-negotiable privacy rules

1. Local-first and private-only are permanent. Never publish populated stores, sources, retrieved memories, logs, credentials, private configuration, or a generally shareable memory link; never expose a populated MCP server publicly. Generic empty source code/specifications may be published only after a privacy scan and explicit approval.
2. Prefer `stdio`. Optional HTTP is disabled by default and may bind only to `127.0.0.1` (unless another private boundary is explicitly approved), with local authentication/session protection, validation, request/rate limits, and a kill switch.
3. Remote access requires a verified account/workspace-scoped private connection. Before the first transmission, verify target, permissions, transport, and current data-use/retention terms. A tunnel secures reachability, not retention policy. Disclose that the request and selected returned records leave the local process for model processing.
4. Never store passwords, keys, tokens, authentication codes, banking data, or other credentials. Add secret/privacy scans, least-privilege files, audit logs, backup/restore, and clean uninstall. Keep stores/backups out of consumer cloud sync unless accepted; encrypt backups where feasible.
5. Use fictional tests in a physically separate store. Return only relevant records, never the complete database by default.
6. Enforce `access_scope` (`local_only`, `remote_allowed`, `workspace_allowed`) plus an authorized principal/account/workspace ACL on every tool path, including search, mutation, counts/status, export, graph traversal, errors, and audit views. A remote surface must never receive `local_only` content.
7. Archive is recoverable, not deletion. Offer a separately confirmed best-effort purge covering selected database content, SQLite WAL/SHM, generated exports, and known backups; retain only a content-free tombstone if required. Report inaccessible/residual copies and explain that local purge cannot retract records previously sent to a remote processor.

### C. Cognitive architecture

Build the following layers:

1. **Working:** bounded request, conversation state, and retrieved packet.
2. **Episodic:** dated conversations, events, and milestones.
3. **Semantic/domain:** durable facts, preferences, decisions, constraints, and consolidated user-specific regions.
4. **Procedural:** reusable rules, methods, and workflows.
5. **Prospective:** future intentions/actions with validity dates.
6. **Temporal/version:** event, learning, validity, change, and confirmation times.
7. **Relationship:** entities and typed links forming a personal graph.
8. **Unresolved:** incomplete projects, questions, blockers, failures, and next actions.
9. **Archive/views:** optional immutable sources plus generated non-canonical Markdown/text views.
10. **Meta-memory/Router:** maps regions and allocates context.
11. **Writer/Consolidator:** adds, corrects, archives, deduplicates, links, and re-ranks.

Support a **shadow mode** that retrieves and scores without injecting memories into the answer. Add a **context governor** that suppresses unnecessary retrieval, enforces local record/token/call/latency budgets, and monitors or stops remote usage only where the platform exposes usable controls.

Native ChatGPT memory, when available in the current ChatGPT surface, remains a complementary recall layer. Do not treat it as interchangeable with the external store; ChatGPT web and local Codex memory may be separate systems.

### D. Canonical data model

Use a local SQLite database as the source of truth. Add migrations and indexes. At minimum, each memory record must support:

- `id`
- `user_scope`
- `access_scope`: local_only, remote_allowed, or workspace_allowed
- `authorized_principal_id` and `authorized_workspace_id` where applicable
- `domain`
- `type`
- `statement`
- `status`: current, historical, partial, disputed, superseded, or archived
- `confidence`
- `importance`
- `source`
- `source_reference`
- `evidence`
- `content_hash`
- `explicitness`: explicit, inferred, or consolidated
- `tags`
- `created_at`
- `updated_at`
- `event_time`
- `learned_at`
- `observed_at`
- `valid_from`
- `valid_until`
- `last_confirmed_at`
- `supersedes_id`
- `archived_reason`
- `sensitivity`
- `retrieval_count`
- `last_retrieved_at`

Store evidence as a minimal source pointer and integrity hash whenever possible, not a duplicate conversation. `event_time` is when the event occurred; `observed_at` is when the source was accessed; `learned_at` is when the memory entered the store. For one user, use one fixed `user_scope`; for multiple users, authenticate and enforce user/ACL scope everywhere—otherwise do not claim multi-user isolation.

Create separate normalized structures for:

- conversations/source documents;
- entities;
- memory-to-entity links;
- typed memory relationships;
- projects and project states;
- decisions;
- unresolved items;
- append-only lifecycle events;
- retrieval feedback and benchmark measurements.

Suggested memory types include:

- FACT
- PREFERENCE
- DECISION
- PROJECT
- PROJECT_STATE
- GOAL
- CONSTRAINT
- DISCOVERY
- WORKFLOW
- SUCCESSFUL_APPROACH
- FAILED_APPROACH
- OPEN_QUESTION
- NEXT_ACTION
- USER_CORRECTION

### E. Source ingestion

Use sources in this order, only when genuinely accessible:

1. explicit statements in the current conversation;
2. native memory/history exposed by the current surface;
3. existing local Codex memories or authorized local files;
4. user-supplied conversations and documents;
5. optional ChatGPT data export.

When importing a large archive, require explicit source scope and size limits, preserve the original files unchanged, and begin with a dry run. Perform distinct passes:

1. inventory conversations, dates, domains, projects, entities, decisions, and outcomes;
2. extract granular candidates with multiple domain probabilities;
3. deduplicate and detect contradictions/temporal changes;
4. link entities, sources, projects, decisions, unresolved items, and next actions;
5. assign provenance, confidence, importance, sensitivity, and validity, then review low-confidence/sensitive inferences.

The review report must show proposed additions, corrections, exclusions, sensitivity flags, and estimated storage/context impact. Commit only approved scope in a transaction, keep an import manifest, and provide rollback for the entire import batch.

Never turn a guess or unstable external fact into a durable personal memory. Direct user statements outrank inferred summaries, and newer explicit corrections outrank older statements.

### F. Domain discovery and brain-region router

Do not impose a universal domain list. Discover domains from the user's actual history and allow them to evolve.

For every domain, store its description, keywords/synonyms, optional semantic representation, related domains, exclusions/overlap rules, activation priority, retrieval depth, and token allocation.

For each request, the router must:

1. decide whether memory helps;
2. detect all relevant domains and useful adjacent relationships;
3. keep unrelated regions dormant and choose shallow/normal/deep retrieval;
4. compile a bounded packet citing memory IDs/sources;
5. label uncertainty/conflict or omit it;
6. never expose unrelated context.

In shadow mode, perform retrieval and ranking entirely inside the local process. Do not return memory text in an MCP/tool response and do not inject it into model context. Store only local memory IDs, scores, latency, and sanitized aggregate measurements unless the user explicitly authorizes more.

### G. Hybrid retrieval and context budgeting

Combine, when locally feasible:

- full-text/keyword search;
- semantic/vector search;
- structured SQL filters;
- graph/relationship traversal;
- temporal queries;
- project, decision, preference, and unresolved-item views.

Rank candidates using a documented, tunable combination of:

- query relevance;
- lexical and semantic similarity;
- domain activation;
- relationship strength;
- lifecycle status;
- confidence;
- importance;
- recency and temporal validity;
- explicit-user-statement priority;
- prior retrieval feedback.

Do not present initial weights as scientific facts. Treat them as hypotheses and tune them against evaluation cases.

Enforce configurable limits for:

- maximum retrieved records;
- maximum context tokens/characters;
- per-domain allocation;
- maximum relationship depth;
- maximum tool calls;
- maximum latency;
- remote-stop threshold where controls exist, otherwise a measured usage alert.

The context governor must skip retrieval for self-contained questions that do not benefit from personal history. Add safe caching where it reduces repeated retrieval overhead without weakening privacy or temporal correctness.

When the budget is exceeded, compress or drop the lowest-value memories while retaining provenance. Never silently truncate the most authoritative current constraint.

### H. Temporal memory and activation decay

The system must answer current-state and historical questions differently.

Required behavior:

1. Newer explicit statements may supersede older ones, but never erase their historical evidence.
2. Corrections create a new record and a version link.
3. `valid_from` and `valid_until` determine when a statement applies.
4. Recency decay changes retrieval priority; it does not silently delete history.
5. Explicit high-importance durable preferences resist decay.
6. Frequently confirmed memories may gain activation strength without becoming more certain than their evidence supports.
7. Dormant or cold memories remain searchable through deep retrieval.
8. Archived memories are excluded from ordinary and deep retrieval; only explicit archive-management or restore operations may access them.
9. Every state change is recorded in the append-only event log.

Track `event_time`, `learned_at`, `valid_from`, `valid_until`, and `last_confirmed_at` separately so the system can distinguish when something happened, when the assistant learned it, when it was true, and when it was last verified.

Implement hot, warm, and cold activation tiers:

- **Hot:** active session, current project state, immediate constraints.
- **Warm:** durable or recently useful domain memories.
- **Cold:** historical, rarely used, or superseded material available for deliberate deep search; archived material is outside activation tiers.

### I. Automatic Memory Writer and consolidation

After meaningful conversations, run or propose a writer pass that classifies each candidate as:

- `ADD`
- `CORRECT/SUPERSEDE`
- `ARCHIVE`
- `NO_CHANGE`

Respect the onboarding mode: make no durable changes in `retrieval-only`; stage every candidate for approval in `review-before-write`; and in `automatic-write-within-approved-scope`, write only non-sensitive, high-confidence records inside the user's approved domains and access scopes.

Before mutation, search duplicates; inspect conflicts/history; mark explicit versus inferred; reject credentials, secrets, chatter, temporary moods, and unsupported conclusions; assign all metadata/links; require approval for sensitive or ambiguous additions; and log the decision/evidence.

Periodically merge redundant wording without losing evidence, create summaries linked to granular sources, promote recurring constraints, lower stale activation, flag unresolved conflicts, update domains/links, and refresh views. Never rewrite the immutable archive.

Controlled consolidation must require adequate evidence—or user approval—before promoting repeated episodes or inferred summaries into durable semantic facts.

### J. MCP operations

Implement and test these required operations:

1. `memory_search`
2. `memory_get`
3. `memory_list`
4. `memory_status`
5. `memory_add`
6. `memory_correct`
7. `memory_archive`
8. `memory_restore`
9. `memory_export_markdown`

Operations 10–19 are optional for the locally functional core and required for locally complete advanced status:

10. `memory_related`
11. `memory_timeline`
12. `memory_get_project_history`
13. `memory_get_decisions`
14. `memory_get_unresolved`
15. `memory_search_conversations`
16. `memory_import`
17. `memory_backup`
18. `memory_record_feedback`
19. `memory_benchmark`
20. `memory_purge` (mandatory maintenance workflow; separately confirmed and never bundled with archive; expose as an MCP tool only if its confirmation boundary is reliable)

Annotate read-only and mutating tools correctly. Validate schemas, enforce ACLs/limits, and return provenance. `memory_export_markdown` and `memory_search_conversations` must be local-only and filtered by default; never return bulk exports or raw transcripts to a remote caller.

### K. Surface integration

#### Codex

- For local Codex, package local skills/plugin, register the `stdio` MCP server, reload, and prove real operations from a new task.
- Cloud Codex cannot use the user's local `stdio` process directly; require a verified private remote path or report it blocked/unsupported.

#### ChatGPT and ChatGPT Work

- Treat integration as optional until the account and product capabilities are verified.
- Use only a private, account/workspace-scoped connection.
- If Secure MCP Tunnel is available, connect the local MCP server through the private outbound tunnel.
- Do not create a public replacement endpoint.
- Verify the current account/workspace data controls, retention terms, administrator visibility, and tool permissions; do not infer them from tunnel availability.
- Explain that access may stop when the user's computer, MCP server, or tunnel client is off.
- Obtain action-time approval before creating a runtime API key or persistent startup service.
- Test retrieval and mutation from the actual target surface before calling it functional.

### L. Cost, quota, and kill switch

Track three different overhead categories rather than combining them: (a) ChatGPT/Codex subscription quota or rate limits, (b) separately billed API/model/embedding/tool usage, and (c) local compute, storage, and electricity. Before paid or quota-consuming steps:

1. inspect current billing/credit state when authorized;
2. identify which actions may consume ChatGPT/Codex quota or API tokens;
3. keep automatic recharge disabled unless explicitly requested;
4. recommend the lowest practical budgets, rate limits, retrieval depth, and context caps, and configure them only when authorized;
5. provide a one-command or one-setting kill switch that stops all remote calls;
6. never imply that the tunnel, API, model calls, embeddings, or hosting are free unless current official documentation explicitly confirms it.

Offer a local lexical/SQLite-only mode that does not require embeddings or API calls. Make semantic embeddings optional and document whether they run locally or remotely.

### M. Reproducible evaluation

Build an evaluation harness comparing:

1. the same surface with the external store absent and native/local memory disabled where possible;
2. the same surface with its available native/local memory but no external store;
3. the same surface with native/local memory plus the extended store.

Run each comparison separately on every advertised surface. Freeze the model/version when selectable, settings, system instructions, prompts, fictional dataset, expected answers, and retrieval configuration. Use isolated test stores/sessions, randomize or counterbalance question order, and prevent memories created in one condition from contaminating another. Include current facts, historical changes, contradictions, cross-domain relationships, unfinished projects, irrelevant distractors, privacy traps, and deliberately unknown answers.

Measure relevant recall; constraint compliance; temporal/conflict handling; unfinished-project and cross-domain quality; irrelevant/hallucinated-memory rates; repetitions avoided; labeled retrieval precision/recall; latency; calls; retrieved records/tokens; measurable incremental subscription/API overhead; and privacy leakage.

Report both absolute results and quality improvement per unit of added overhead. Do not claim an advantage from synthetic or deterministic tests alone.

### N. Build phases

1. **Diagnostic:** detect the real environment; select safe local directories; preserve existing files; establish privacy, authorization, cost, and completion boundaries.
2. **Local core:** create schema/migrations, event log, fictional fixtures, backups, exports, MCP server, and the nine required tools.
3. **Retrieval brain:** add discovered domains, multi-domain routing, hybrid search, budgets, temporal queries, graph links, and unresolved-project views.
4. **Writer brain:** add candidate extraction, deduplication, contradiction detection, versioning, secret rejection, sensitivity rules, and consolidation.
5. **Import:** dry-run only authorized sources, preserve originals, review uncertain candidates, and support transactional rollback.
6. **Codex:** install locally, reload, and prove real MCP calls from a new task.
7. **Optional ChatGPT/ChatGPT Work:** recheck official capabilities and policies, obtain action-time approval for credentials/persistence, privately connect, and test end to end.
8. **Benchmark/hardening:** start in shadow mode; measure quality, overhead, and privacy; add recovery, health checks, uninstall, and failure documentation.

Pin dependencies and exact runtime versions. Reproduce installation from a clean copy in a fresh environment, not only the development tree.

### O. Required tests

With fictional data, verify every applicable implemented or claimed feature: add/retrieve; exact duplicates (plus semantic duplicates for advanced); correction/history; current versus historical answers; archive/restore; secret rejection; domain routing; advanced graph/unresolved behavior; exposed budgets; private binding; exports; backup/restore; clean install; real calls on each claimed surface; kill switch; safe uninstall; and separately confirmed purge across database/WAL, exports, and known backups, reporting unavoidable residual-storage limits.

### P. Completion criteria

Report three levels separately:

- **Locally functional core:** reproducible clean install; SQLite migrations; nine required operations passing automated and real MCP tests; lexical retrieval, routing, temporal correction/archive, secret/privacy/budget tests, event log, export, backup/restore, separately confirmed purge, and uninstall verified.
- **Locally complete advanced:** the core plus operations 10–19, semantic retrieval, graph traversal, writer/consolidation, advanced temporal/unresolved behavior, and controlled local benchmark.
- **Cross-surface complete:** the locally complete advanced system plus real scoped retrieval and mutation tests from ChatGPT, ChatGPT Work, and Codex. If any surface is unavailable, say `blocked` or `unsupported`; do not call the cross-surface project complete.

Never promote any level because static files merely exist. Disclose unavailable history, unsupported surfaces, untested optional features, actual/potential costs, permissions, persistent services, and behavior while the computer is off.

### Q. Final report

Provide a concise component table with status (`complete`, `locally functional`, `partially connected`, `untested`, `blocked`, or `unsupported`), verification evidence, actual/potential cost, permissions/persistence, smallest remaining action, and offline behavior. Also report paths, memory counts, tests, genuinely working surfaces, retrieval/context budgets, backup/restore/uninstall commands, benchmark results, limitations, and unverified assumptions.

Begin now with the diagnostic. Build everything that can be completed safely, locally, privately, and without paid services. Stop only at a genuine authorization, credential, cost, or capability boundary. Never substitute confident language for missing verification.

— END OF REPRODUCTION PROMPT —

Final note

If you test this, please share results—not private memory content. The most valuable contribution would be a reproducible comparison showing whether the extra recall and continuity are worth the additional context, quota, latency, setup, and maintenance.

Sorry I’m new to this dev board. However far from new. I like this idea, how do I connect w you? I activated my dev account just to connect w you after swearing I never would.

Shaun

Haha, I’m honored that this project made you do something you swore you’d never do :laughing::sweat_smile::partying_face::smiling_face_with_sunglasses:

…and to answer your question about connecting, it depends on what your goal is. If you have questions that might benefit everyone interested in the project, you can post them here. If you want to chat with me privately, you can contact me via the “My messages” menu on the left. If you mean something else entirely, just let me know and I’ll get back to you.

Thanks,
Sonic-The-Hedgehog

I’m coming at this from a very different direction. I’m not a developer — I’m a sheep farmer and contractor who became a very heavy ChatGPT user.

I kept running into the problem of having useful AI conversations and decisions disappear into chat history. I also didn’t want to become the AI’s “memory courier,” constantly deciding what to save and feeding it back.

So my AI (I call him TARS) and I gradually built a simple external system we call Portable TARS: a small “Boot Box” for how TARS operates, a “Port Box” for durable context, an “Active World” for what’s happening now, and Decision History that preserves important decisions and why we made them.

What caught my attention in your post is how many of the same rules we arrived at independently — targeted retrieval, current vs. historical information, explicit statements vs. AI inference, superseding stale information, and having the AI help maintain its own continuity.

Ours is nowhere near as technically sophisticated as yours, but we’ve been using it every day across farming, construction, real estate, finances, family, etc., and learning from what breaks.

I wrote up what we’ve learned in an article called “We Stopped Using ChatGPT Like a Chatbot.” I’m happy to share it once the forum decides I’m trustworthy enough to post links. :slightly_smiling_face:

I’d be interested in comparing notes. You’re attacking this from the technical side; we’ve mostly been attacking it from the “use the damn thing every day and see what breaks” side.