Back to all writing

I Wanted to Build My Own JARVIS

SYNTHESIS Foundation•Part 1•8 min read•By Mikhil

SYNTHESIS began with an embarrassingly ambitious idea:

I wanted to build my own JARVIS.

Not a chatbot wearing a futuristic interface. Not a voice assistant that could set a timer, report the weather, and misunderstand everything else.

I wanted an assistant that could understand what was happening around it, remember context, use tools, complete tasks, and communicate naturally.

Something that felt less like opening an application and more like working with an intelligent presence.

I knew that was an enormous idea.

I wanted to attempt it anyway.

The Marvel inspiration

The original inspiration came from several Marvel artificial intelligences rather than one character alone.

JARVIS represented reliability, competence, and the feeling that the assistant understood the larger situation—not only the latest command.

FRIDAY represented a more direct, operational intelligence: an assistant capable of providing information while events were happening.

Vision represented reasoning, restraint, and a form of intelligence that was more than a collection of automated responses.

Ultron represented the opposite lesson: intelligence combined with power, autonomy, and the wrong interpretation of a goal.

I was fascinated by what might happen if those qualities were considered together.

I did not want to copy any of those characters literally. Their personalities, fictional capabilities, and worlds belong to stories.

But they gave me a way to think about the qualities an advanced assistant might need:

  • Awareness of context

  • Reliable action

  • Memory across interactions

  • Independent reasoning

  • A recognisable identity

  • The ability to use other systems

  • Restraint when power creates risk

That combination eventually suggested the name SYNTHESIS.

The project was not meant to reproduce one fictional assistant. It was an attempt to synthesise several ideas about what an assistant could become.

Chatting was not enough

Modern AI systems can hold impressive conversations.

They can explain concepts, generate plans, write code, summarise information, and respond in remarkably natural language.

But conversation alone was not what I wanted.

If I asked an assistant what was happening on my computer, it should be able to observe relevant system context rather than guess.

If I asked it to complete a task, it should be able to identify the necessary tools, use them in the correct order, and determine whether the task actually succeeded.

If something changed while it was working, it should be able to respond.

If we had already discussed an important preference or decision, I did not want every conversation to begin from zero.

The difference was action.

A chatbot primarily produces answers. The assistant I imagined needed to connect language to perception, memory, decisions, and tools.

That made the project far more complicated than adding a voice to an AI model.

The first idea was dangerously broad

My earliest ambition was that SYNTHESIS should eventually be able to do almost anything I asked, provided we could build the required tools.

That idea was exciting because it removed artificial boundaries from the imagination.

The same assistant might one day understand my screen, manage applications, organise information, conduct research, help with projects, monitor ongoing work, and communicate through different interfaces.

But “do almost anything” immediately creates another question:

Should it do everything it is capable of doing?

An assistant able to act on a computer can also make mistakes on that computer.

Opening an application is different from deleting information. Reading system status is different from sending something in the user’s name. Suggesting a command is different from executing it.

Capability without boundaries would not create a better assistant.

It would create a more dangerous one.

That realisation did not shrink the ambition behind SYNTHESIS. It changed the definition of success.

The goal could not only be to make the assistant powerful.

It also needed to make that power understandable and controllable.

Fiction skips the permission screen

In films, an intelligent assistant often performs complex actions immediately.

That makes sense for storytelling. Nobody wants to watch the hero review a confirmation dialogue in the middle of an action sequence.

Real systems do not have that luxury.

An AI can misunderstand an instruction. It can choose the wrong tool, supply an incorrect argument, or report success when an action silently failed.

Even a correct instruction may have consequences the user did not fully intend.

That means a real assistant needs friction in the right places.

Harmless actions may be allowed to happen automatically. More consequential actions require greater care. Dangerous or irreversible actions should require explicit confirmation.

The assistant also needs to verify outcomes rather than assuming that calling a tool means the task was completed.

This principle became central very early:

As an assistant becomes more capable, it must also become more accountable.

SYNTHESIS should not hide important actions behind the illusion of intelligence. It should make clear what it intends to do, recognise when permission is required, and honestly report what happened.

I did not know how to build it

There was another obvious problem.

I was not a programmer capable of independently engineering an advanced agent from an empty folder.

I understood the experience I wanted. I could describe behaviours, evaluate ideas, question unsafe decisions, and decide what the system should become.

But my implementation ability did not match the scale of the project.

AI made the attempt possible.

I could turn the large vision into smaller requirements, discuss architecture, generate code, run tests, show errors, challenge weak approaches, and continue iterating.

That did not suddenly make me an expert.

I could still misunderstand generated code. I could accept something because it sounded technically convincing. I could create a demonstration that appeared impressive while remaining unreliable underneath.

So SYNTHESIS became two experiments at once.

The first was about building a personal AI assistant.

The second was about how far someone without conventional programming mastery could direct the construction of a complex system using AI.

The project needed layers

The fictional idea initially existed as one enormous sentence:

Build an assistant that understands everything and can do anything.

That is not an actionable plan.

The idea became more realistic when I started separating it into layers.

Perception would help the assistant understand relevant parts of its environment.

Reasoning and agency would help it break goals into steps and decide what to do next.

Tools and orchestration would connect decisions to real actions.

Memory would preserve useful context across interactions.

Identity and interface would determine how communicating with it actually felt.

Permissions and verification would constrain its power and check whether actions succeeded.

These layers transformed SYNTHESIS from a fictional wish into a project that could be approached one piece at a time.

The complete vision remained distant.

But the next test became possible.

Utility before spectacle

A project inspired by science fiction naturally creates an urge to begin with the most cinematic parts.

A glowing interface. A distinctive voice. A wake word. Animated visualisations. The feeling of talking to something futuristic.

I still want SYNTHESIS to eventually feel special.

But an assistant with a beautiful interface and no dependable abilities is only a demonstration.

The practical foundation matters first.

Can it hold a useful conversation?

Can it understand its available tools?

Can it perform a harmless action correctly?

Can it refuse to perform a dangerous action without confirmation?

Can it determine whether the action actually worked?

Can it preserve the context necessary for the next step?

Those questions are less cinematic than designing a JARVIS-style interface.

They are also what separate an assistant from a prop.

The development direction therefore became laptop-first and utility-first: make the most useful and dependable system possible with the hardware available, then consider heavier capabilities only when real limitations appear.

An ambition with boundaries

SYNTHESIS is still far from the assistant I originally imagined.

It does not understand everything happening around me. It cannot safely perform every task. It does not possess fictional intelligence, and I do not want to pretend that a few successful tests make it autonomous in the way people imagine from films.

But the original ambition still matters.

It establishes the direction: an assistant that can understand context, communicate naturally, use tools, remember what matters, and become increasingly useful without quietly taking control away from its user.

JARVIS made the idea exciting.

Ultron made the danger obvious.

Building SYNTHESIS means trying to preserve the ambition while respecting the warning.

The question is no longer merely:

Can I build an assistant that does things?

It is:

Can I build one that deserves to be trusted with the ability to act?