Back to all writing

A Chatbot Wasn’t Enough

SYNTHESIS Foundation•Part 2•8 min read•By Mikhil

The first version of my SYNTHESIS idea was easy to describe:

I wanted an AI assistant I could talk to naturally.

But the more I thought about that sentence, the less complete it felt.

I could already talk to AI systems. They could answer questions, explain concepts, generate plans, and help me write code.

That was useful, but it was not the assistant I had imagined.

If SYNTHESIS could only reply to messages inside a text box, then it would still be separated from the world in which I wanted its help.

It might tell me how to complete a task.

It could not complete the task.

It might explain what could be happening on my computer.

It could not understand the current situation directly.

It might produce a good plan.

It could not follow that plan, react when something failed, or verify that the intended outcome was achieved.

A chatbot was part of SYNTHESIS.

It could not be the whole system.

Language is only one layer

Large language models create the visible intelligence in many AI products.

They interpret instructions, reason through problems, and generate responses. Because conversation is the part users see, it is easy to treat the model as the entire assistant.

But a model alone does not automatically know what is happening outside its conversation.

It does not inherently understand my current screen, available applications, system state, files, ongoing work, or previous decisions. It cannot safely use tools merely because it can describe how those tools should be used.

A dependable personal agent would require several systems working together:

  • A language model to understand and reason

  • Perception to gather relevant context

  • Memory to preserve useful information

  • Tools to perform real actions

  • Permissions to control those actions

  • Execution logic to manage the process

  • Verification to check the result

  • An interface through which all of this feels understandable

The assistant was not one magical piece of software.

It was a coordination problem.

Perception grounds the conversation

Suppose I ask:

What is happening on my computer?

A chatbot without perception has very little to work with. It might ask me to describe the situation or provide general troubleshooting advice.

SYNTHESIS should eventually be able to gather relevant context itself.

That could include information from the system, the current screen, a microphone when explicitly active, or other approved sources.

But perception is not simply “let the AI see everything.”

Context should be relevant, permission-based, and limited to what the task requires. An assistant should not constantly collect information merely because that information might become useful later.

Giving a system access to context creates both capability and responsibility.

It needs to know when observation is appropriate, what it is allowed to inspect, and what should remain private.

Perception should ground the assistant’s reasoning.

It should not become an excuse for invisible monitoring.

Memory creates continuity

Without memory, every interaction begins again.

The assistant may understand the current message perfectly while forgetting decisions made five minutes, five days, or five projects earlier.

A personal assistant should be capable of continuity.

It should remember useful preferences, unfinished tasks, prior decisions, and relevant context. If I have already explained how I want something handled, I should not need to repeat the entire explanation every time.

But unlimited memory would create its own problems.

Not everything deserves to be remembered. Some information becomes outdated. Some should expire. Some should never be stored.

A system that remembers carelessly can become as frustrating as one that forgets everything.

Memory therefore cannot be treated as a box where every conversation is permanently collected.

It needs judgment:

What is worth preserving?

How long should it remain relevant?

Can the user inspect, correct, or remove it?

SYNTHESIS needs continuity without treating my life as raw material to store forever.

Tools turn language into action

Conversation becomes agency when the system can use tools.

A tool might retrieve information, inspect approved system status, interact with an application, organise something, or perform another defined action.

The model does not perform these actions directly. It decides when a tool may be useful and supplies the information that tool requires.

That separation matters.

A model can generate convincing language even when it is wrong. A tool operates on the real world, where mistakes have consequences.

Therefore, a tool needs a clear purpose, controlled inputs, an assigned level of risk, and a way to report what actually happened.

SYNTHESIS should not claim success simply because it attempted an action.

It needs evidence that the action succeeded.

This distinction between execution and verification became one of the foundations of the project.

A request may require several steps

Real tasks are rarely one action long.

A useful assistant might need to understand a goal, identify the available tools, break the goal into smaller steps, perform them in sequence, and adjust when the environment responds unexpectedly.

That is where the idea of an agent becomes important.

An agent does more than generate the next sentence. It maintains progress toward an outcome.

But multi-step execution also creates more opportunities for failure.

An early misunderstanding can affect every later step. A tool might return incomplete information. The environment might change. A plan that looked sensible before execution might stop making sense halfway through.

The assistant needs to observe results and reconsider what to do next rather than blindly completing a prewritten sequence.

Autonomy should not mean stubbornness.

Orchestration is the hidden work

If SYNTHESIS eventually interacts with multiple tools, applications, models, and sources of context, something has to coordinate them.

Which tool is appropriate?

What information should be passed to it?

Does the action require confirmation?

What happens if the tool is unavailable?

Should the assistant retry, choose another approach, or stop?

What context should be carried into the next step?

This coordination is orchestration.

It is less visible than a futuristic interface, but it determines whether the assistant behaves like a system or a collection of disconnected tricks.

A demo can call one tool successfully.

An assistant needs to manage uncertainty across many possible actions without confusing capability with permission.

Identity is more than a voice

I also wanted SYNTHESIS to feel like one coherent assistant.

That does not mean giving it a dramatic personality and pretending it is conscious.

Identity can be practical.

It includes how the assistant communicates uncertainty, asks permission, reports failure, remembers preferences, and behaves consistently across different interfaces.

A system that sounds friendly in conversation but becomes vague when something fails does not have a trustworthy identity.

Neither does one that speaks confidently when it lacks evidence.

The voice, visual interface, language style, and interaction patterns should eventually support the same underlying character: capable, direct, context-aware, and honest about its limitations.

That identity must emerge from behaviour, not only appearance.

The terminal is a workshop

When I first started seeing SYNTHESIS run inside a terminal, I wondered whether that was where it would always live.

It will not.

The terminal is useful because it exposes what the system is doing. During development, I can see registered tools, permission decisions, execution results, failures, and diagnostic information without hiding them behind a polished interface.

It is a workshop.

It allows the internal pieces to become dependable before I build the room in which people will eventually experience them.

Later, SYNTHESIS may communicate through text, voice, shortcuts, a tray interface, or other forms. Those decisions matter, but they should sit on top of a working system.

A beautiful interface cannot compensate for unreliable execution.

Utility before spectacle

The most tempting version of SYNTHESIS begins with the parts that resemble science fiction.

A wake word. A cinematic interface. Natural voice. Animated graphics. The appearance of intelligence.

The more useful development order begins elsewhere.

Can the assistant understand an ordinary request?

Can it use practical tools reliably?

Can it fetch current information when required?

Can it manage applications or system functions within safe boundaries?

Can it remember enough context to remain useful?

Can it ask before doing something consequential?

These abilities are less dramatic in isolation.

Together, they create the foundation that makes the dramatic parts meaningful.

The plan became to build the strongest practical system possible on my existing laptop, prioritising utility first. Heavier computing or paid services should become relevant only when a real capability reaches a genuine hardware limit.

From chatbot to agent

SYNTHESIS is not defined by one model.

It is the combination of language, perception, reasoning, memory, tools, permissions, execution, verification, and interface.

Each layer introduces new capability.

Each also creates new ways for the system to misunderstand, overreach, or fail.

That is why developing SYNTHESIS cannot be reduced to finding the smartest model and connecting it to my computer.

The harder work is designing how the pieces cooperate—and where they must stop.

A chatbot answers:

Here is what you could do.

The assistant I am trying to build should eventually be able to say:

I understand the situation. Here is what I can do, here is what requires your permission, and here is what actually happened.