AI agents

An AI agent is not a chat bubble glued to the corner of the screen. It is a layer that reads your documentation, reaches into your API and finishes the task - with the source of every answer visible and a confirmation before any operation that cannot be undone.

We have worked with leading companies and startups

Challenges

Three moments when an agent stops being an experiment

Not sure whether your process is a fit for an agent at all? Book a consultation - we will say plainly when a plain form or better search solves it more cheaply and more reliably.

  • „Support answers the same questions over and over”

    The first line burns a working day on questions whose answers sit in the terms of service and in past tickets. The agent takes over the repetitive traffic, cites the document it used, and hands everything it does not understand to a human together with the conversation context - without asking the user to describe the case a second time.

  • „Knowledge sits in documents nobody opens”

    Manuals, policies and years of decisions live across several systems at once, and finding the right paragraph takes a quarter of an hour. We index those sources and build search that answers in a sentence and shows which document it came from. An answer without a source never reaches the user.

  • „A simple task costs the user five screens”

    Changing a date, correcting details, generating a report - trivial actions spread across a form, a filter and two confirmations. The agent finishes them with a single instruction in natural language, calling exactly the endpoints you already have. We do not write new business logic, we expose the existing one.

Who it is for

Who we deploy agents for

  • Product teams handed "add AI" with no stated goal

    The decision came from the board, the deadline exists, the scope does not - what remains is pressure and the risk of building a feature nobody uses. We start by sorting processes into those an agent will genuinely shorten and those better left alone. The workshop ends with one process for the pilot and the arguments you need to defend that choice internally.

  • Companies with knowledge scattered across systems

    Terms in one place, decisions in inboxes, procedures in an intranet nobody browses. We build a search layer that reaches those sources and cites the specific passage, and alongside it a maintenance process - because a knowledge base without an owner goes stale within months.

  • Products with customer support on the front line

    Ticket volume grows faster than the team and hiring cannot keep up. The agent takes over the repetitive threads and prepares the rest for handover - with a summary, a history and a suggested reply for the consultant. The goal is not to cut the user off from a human, but to shorten the path to one.

  • Teams that already have a chatbot and want it to finally work

    The bot sits on the site, users click "connect me to an agent" in their second message, and the statistics look good only in the vendor's report. We go through the conversation log, point out where the dialogue falls apart, and rebuild the script and the knowledge sources instead of swapping the model for a newer one.

Four printed interface layouts on a desk: a dashboard, a phone screen, a form and a table, with one block circled in marker

Our difference

What makes our approach different

  • We start from how the agent gets it wrong

    Most projects design the happy path and bolt error handling on at the end. We start from a catalogue of failures - misread intent, missing data, contradictory sources, an action applied too broadly - and design the interface behaviour for each one. Only then does the version where everything goes right get built.

  • The agent acts, it does not only answer

    An assistant that merely summarises documents ends up as a better search box. We connect the agent to your endpoints so it closes the task inside the same conversation - with an explicit list of operations it is allowed to perform and a threshold above which it asks the user for consent.

  • Your expert writes the control questions

    The set we measure quality against is created on the client side - your specialist is the one who knows which answer is correct and which merely sounds credible. We supply the method and the tooling, the judgements stay yours. That way, once we leave, you still have something to measure every future change against.

  • The pilot has a date and an exit gate

    After three weeks there is a verdict, not another request for an extension. We agree the criteria before the start so the decision does not hinge on how well the demo went. A "not worth it" outcome counts as an outcome - with a report on what specifically failed and under what condition the topic returns.

Process

From workshop to production in 8–10 weeks, with an exit gate after the pilot

  1. Week 1

    Scope and feasibility workshop

    We break the candidate processes down into factors: how many people use them, how expensive a mistake is, whether a single source of truth exists. Some ideas are dropped here deliberately - because a rule solves them, not a model. We leave with one process for the pilot and a written justification of the choice.

  2. Week 1–2

    Data audit and the control set of questions

    We check what the agent will feed on - how current the documents are, duplicates, contradictory versions, sensitive data to cut off. In parallel your experts write control questions together with model answers. Those, not our impression, are what will measure quality later.

  3. Week 2–3

    Conversation prototype and edge states

    A clickable prototype showing not only the happy path but also refusal, misread intent, an interrupted action and handover to a consultant. We test it with a few people from your team before any backend exists - fixing the conversation script then costs hours rather than a sprint.

  4. Week 3–5

    Pilot with a narrow group

    The agent goes to a limited group of users with conversation logging switched on. We measure accuracy against the control set, the share of handovers to a human, latency and the cost of a conversation. This stage ends at a gate - we go further, or we close the topic with a report on why it is not worth it.

  5. Week 5–8

    Production rollout with guardrails

    The final interface in your design system, usage limits, filtering of personal data before it reaches the model, handling of provider outages and a kill switch on your side. The release ships behind a flag, so you widen the audience gradually rather than jumping to the whole user base.

  6. Week 8–10

    Measurement, tuning and hand-off

    We tune prompts, tool descriptions and the index against real conversations, not workshop assumptions. The team receives documentation, access to the log and a procedure for adding the next task to the agent - along with instructions for extending the control set, so a new capability does not break an old one.

Deliverables

Three layers. The model is the smallest of them

  • A map of tasks, tools and boundaries

    A list of the tasks the agent takes over, each with a tool assigned on your side - an API endpoint, a knowledge-base lookup, an action in the panel. Every task carries a decision threshold: what the agent does on its own, what it hands to the user for approval, what it never touches under any scenario. Plus a risk register in which every entry has a braking mechanism attached.

  • The agent's interface inside your product

    A Figma design and components consistent with your design system - entry into the conversation from wherever the user already is, a streamed answer with a link to its source, an honest "I don't know" state, confirmation of irreversible actions, undo and handover to a human. We design the failure states as carefully as the successful ones, because they are what settles trust in the whole feature.

  • The technical layer and the control set

    Model integration, a knowledge index with citations, tool orchestration with a step limit and a fallback path, guardrails on the input and on the output. Plus a control set of questions built with your experts - every release passes through it before deployment, and the result is a number, not an impression from a presentation.

Toolkit

What we apply

Some of this is a public obligation, some are our own design rules - drawn from the fact that an agent makes its mistakes in front of the user, and that moment decides whether anyone comes back to it a second time.

A workbench from above: two printed interface layouts, a magnifier enlarging one block, a steel ruler and a greyscale step wedge running from white to black, with one line underlined in orange marker
  • RAG with a link to the source on every answer
  • Human confirmation before any irreversible action
  • A control set of questions before every release
  • Explicit notice that a system is running the conversation (AI Act)
  • Versioned prompts and tool definitions in the repository
  • A conversation log with cost, latency and answer ratings
  • WCAG 2.2 AA for the conversation interface
  • Data minimisation - the model never sees fields it does not need

Next step

Let us talk about your project

Tell us what you have to do. We will come back with scope and a quote.

Selected work

Work from the products we wire agents into

A selection of projects from the product types where agents come up most often - client panels, web applications and service systems. The AI deployments themselves live in client environments, so we show them during a consultation.

Projects

Outcomes

What you get, on what timeline and how you verify it

  • 3 weeks

    From the workshop to a go/no-go decision after the pilot - the point at which you can walk away holding a measurement rather than a presentation

  • 8–10 weeks

    From the workshop to an agent in production for one process in one channel; further processes are added as separate releases

  • 0 silent actions

    No irreversible operation runs without the user confirming it - that is a design rule, not a switch in the configuration

  • 30 and 90 days

    The points at which we report accuracy against the control set, the share of handovers to a human and the cost of a single conversation

The timelines come from the schedule described above on this page, and WCAG 2.2 AA together with the AI Act disclosure duty are public requirements, not our promise. We do not average client results - conversation logs and control sets live in client environments, so we show them by name during a consultation.

Client voices

What our clients say

  • „Working with the Wzór team was a genuinely great experience! They always answered our questions quickly and were flexible about our comments. They have no shortage of creative ideas and are at the same time thoroughly reliable and on time. The end result fully met our expectations; we recommend them to anyone looking for professionals who care.”
    Maja Wieloch-Silecka Operations Director, Bosfor Group
  • „Our clients expect us to act fast and effectively across a very wide scope - from building a brand, through a website with a store, to running an effective sales campaign. Not an easy task, but doable when you have a tight-knit team and partners who take on challenges on the fly. For us, quality, accountability and commitment are key, and working with UX Agency Wzór gives us that 100%.”
    Piotr Alberski Chief Operating Officer, Agencja XO Media
  • „Our cooperation with UX Agency Wzór started with our own website. We really liked the way they work, so we decided to subcontract projects for our clients to them on a White Label basis. We loved their workflow, the way they hand projects over to our team, their documentation and their work on the design system. Now we can take on bigger risks.”
    Mateusz Swół Chief Operating Officer, Sellace

Time and scope

How long deployment takes and what the quote depends on

8–10 weeks

Eight weeks covers one process in one channel, with data that is already in order. Several processes at once, integrations with systems that have no ready API, or a knowledge base that needs cleaning push the deadline further - we then split the work into releases so the first useful version reaches users sooner.

  • The number of processes and channels the agent has to serve

  • The state of the knowledge base - currency, duplicates, contradictory versions

  • The scope of writing actions and the readiness of your API

  • Requirements around sensitive data, retention and processing location

  • Whether we design the interface too, or the agent enters an existing design system

FAQ

Questions that come up most often

How much does deploying an AI agent cost?

We quote each deployment individually, because two projects under the same name can differ several times over in effort. The quote is decided by: the number of processes and channels, the state of the knowledge base (the more contradictory and outdated documents, the more work happens before the first prompt), the scope of writing actions and the readiness of your API, requirements around sensitive data and retention, and whether we design the interface from scratch or it enters an existing design system. On top of that comes the running model cost, which depends on conversation volume - we show it separately so it does not blend into the cost of the deployment. You get a concrete proposal within 48 hours of a short brief.

How is an agent different from an ordinary chatbot?

A chatbot walks a conversation down a decision tree drawn in advance and stops where the script stops. An agent is given a goal, a set of tools and boundaries - it decides for itself which tool to reach for to finish the task, and it can stop halfway when it lacks data or permissions. The practical difference is that an agent performs operations inside your system rather than only talking about them.

Where does the agent get its knowledge, and can it make things up?

The agent answers from your sources through semantic search (RAG), not from what the model memorised during training. Every answer carries a link to the document it came from - and when it finds nothing that fits, it says "I don't know" and offers a human instead of guessing. The risk cannot be driven to zero, so we measure it openly: every release passes through a control set of questions with model answers written by your experts, and we report the score as a number before each deployment.

Does our data reach the model, and where is it processed?

The model only sees the fields the task requires - we strip the rest before sending, together with any personal data that can be removed from the query. The processing location, retention and a ban on training with your data are settled while choosing the provider and written into the data processing agreement; for teams with a hard EU-processing requirement we pick a provider that meets it. All traffic goes through your backend, so keys and logs stay on your side.

Which model do you use?

The one that wins on your control set at an acceptable cost and latency - we settle that by measurement during the pilot, not by a claim on a product page. The model-calling layer is kept separate, so switching provider is a configuration change rather than a rewrite of the deployment. That matters, because in this category the ranking shifts faster than most products ship a release.

Can the agent perform actions, not just answer?

Yes, and that is usually the reason to deploy one. The agent receives an explicit list of operations it may perform, wired into your existing endpoints - we do not write new business logic. It performs reversible, low-risk operations on its own and shows anything irreversible to the user for confirmation. Every call lands in the log alongside the conversation that preceded it, so after-the-fact audit does not involve guesswork.

We already have a chatbot that does not work - do you start from scratch?

No. We start from the conversation log, because that is where you see the moment users give up and what they were actually looking for. The problem usually sits in the script, in the knowledge sources, or in the bot having no way to hand a case to a human - not in the model. If it turns out the trouble runs wider than the chat itself, we say so plainly and propose a UX audit instead of swapping the model for a newer one.

What if the pilot goes badly?

We close it with a report and a decision, not a request for an extension. The criteria are agreed before the start, so the verdict does not depend on how the demo went. The report states what specifically failed - missing data, a process too volatile to automate, or a conversation cost with no path to paying for itself - and under what condition the topic is worth revisiting. A negative result after three weeks is cheaper than a deployment nobody uses.

Book a free consultation with an expert to discuss your project and get answers to your questions.

- Patryk Korycki, CEO

Schedule a meeting
Book a free consultation