Skip to main content

Frequently asked questions

Straight answers about how we work, what things cost, and how we keep your data safe.

Working with us

How ChainCraft Global engages, what we take on, and what we decline.

We design, build, and operate AI systems for enterprises: autonomous agents, chatbots, voice AI, document automation, workflow orchestration, AI-native apps and SaaS products, integrations into existing systems, and the strategy work around all of it. Every engagement is aimed at a system running in your business, not a report about one.

Every engagement starts with a fixed-fee discovery (typically 4–6 weeks) that produces a committed scope, timeline, price, and expected ROI. You decide with full information. Shorter feasibility spikes of one to three weeks are available where a single technical assumption is the main risk.

Both. Our sweet spot is mid-market and enterprise operations work, plus funded startups building AI products with our AI SaaS practice. What matters more than size is whether a named owner exists on your side who will run the system once it is live.

Yes — rescue engagements are common. We begin with a technical audit and give you an honest keep/fix/rebuild assessment, including the cases where the right answer is to stop. The audit is fixed-fee and useful whether or not you continue with us.

Yes. We serve clients across eight countries with overlapping-hours delivery teams, including the US, UK, Singapore, the UAE and across Europe. Contracting, data residency and working hours are arranged per engagement.

Systems designed to deceive people about whether they are interacting with AI, cold-calling and collections voice agents, applications that make consequential decisions about individuals without human oversight, and anything whose primary purpose is surveillance of employees. We will also decline work we do not believe will reach production.

Yes, happily. Send yours or use our mutual NDA template. We can have one executed before a discovery conversation if that is easier for your legal team.

Mostly mid-market and enterprise organisations (roughly ₹50 crore revenue and above), plus funded startups building AI products. Below that scale, our fixed-fee discovery is often still worthwhile even if the delivery is better done elsewhere, and we will say so.

Yes — 30 minutes with a delivery architect, no sales pitch, and you keep the notes. If the honest answer is that AI will not help with your problem, you will hear that on the call rather than after a proposal.

The architects who scoped it. We staff 3–5 senior engineers with a named lead accountable to you, and we do not move people to a junior bench after the sales process. The person who wrote your proposal is in your weekly demo.

Yes. We deploy inside client VPCs regularly across AWS, Azure and GCP, and align with SOC 2, GDPR, India’s DPDP Act and sector-specific requirements. Where fully private model hosting is required, we deploy open-weight models with no data leaving your environment.

Engagements are milestone-based, so you can stop at any milestone boundary and you keep everything produced to that point — code, prompts, evaluation data and documentation. There are no long lock-ins and no termination penalties.

Delivery & timelines

How long things take, what we need from you, and what happens after launch.

Focused deployments (a chatbot, one automated workflow) ship in 4–8 weeks. Agent systems and full products run 8–16 weeks. Timelines are committed at the end of discovery, not guessed at the start.

Plan for a product owner at roughly 25% time and subject-matter experts for a few hours weekly. We handle the rest, and we train your team as we go. Engagements where nobody on the client side has time are the ones that fail, so we say this early.

Every system launches with monitoring, KPI dashboards, and runbooks. You can own operations, or we run it under a managed agreement with SLAs. Either way, the handover material exists from day one rather than being written at the end.

Yes — tiered support agreements covering monitoring, model updates, evaluation re-runs, and continuous improvement. Support is always optional; nothing we build is engineered to require us.

3–5 senior engineers with a named lead accountable to you. We deliberately avoid large, diluted teams — coordination overhead grows faster than throughput, and AI delivery rewards depth over headcount.

Yes — we run overlapping-hours delivery for clients across India, the Middle East, Europe, the UK, the US, and Singapore. Weekly demos are scheduled in your working hours, not ours.

A live demo on your real data, a written update covering progress, risks and decisions needed, and asynchronous access to the team throughout. We default to writing things down, which makes progress reviewable rather than reported.

Scope is fixed after discovery and changes are handled as explicit, priced change requests rather than absorbed silently. Where a change replaces existing scope of similar size, we swap it at no cost — most changes during delivery are trades rather than additions.

The system runs against real production data, proposing actions without executing them, while humans continue working normally. We compare its proposals against actual human decisions and drive the disagreement rate down before granting any write access. It is the single most effective de-risking technique we use.

Timelines are committed after discovery and we hold them. Where a delay is caused by something on our side, we absorb it. Where it is caused by data access or decisions on your side, we flag it in the weekly update as soon as it appears rather than at the deadline.

Your engineers sit in on delivery from the start, the codebase is reviewed with them, and handover includes runbooks, architecture documentation and working sessions on the parts they will operate. We consider it a failed engagement if your team cannot change the system after we leave.

Sometimes, for a bounded first slice where the requirements are already clear. We will not commit a full scope and price without discovery, because those commitments would be fiction — and a fixed price built on a guess helps nobody.

Security & data

What happens to your data, where it goes, and who can see it.

No. We use enterprise API tiers with zero-retention agreements, and we can add a redaction gateway so sensitive fields never leave your environment. This is contractual with the model providers, not a policy statement — we can share the relevant terms.

Yes — AWS, Azure, and GCP in-VPC deployments are routine, including fully private model hosting where required. For clients who cannot send data outside their environment at all, open-weight models run entirely on your infrastructure.

Our delivery practices are SOC2-aligned, GDPR-aware by default, and we run EU AI Act readiness assessments as part of governance engagements. We map your specific systems to their obligations rather than making a blanket compliance claim.

You do — code, prompts, evaluation suites, fine-tuned weights, dashboards. Everything, from the first commit. There is no licensing arrangement and no component we retain.

Yes — DPAs, NDAs, and sector-specific addenda are standard parts of our contracting. We also provide a subprocessor list naming every model and infrastructure provider in the chain.

Layered defences: input classification, constrained tool permissions, output filtering, and red-team testing before launch — plus monitoring that flags anomalous behaviour in production. Any content the model reads that a user did not author is treated as untrusted data rather than instruction.

A service that sits between your systems and any model provider. It replaces sensitive fields — names, identifiers, account numbers, health data — with reversible tokens before the request leaves your environment, then restores them in the response. For summarisation, classification and drafting the model does not need real values, so quality is unaffected.

Only the named engineers on your engagement, under role-based access, with access logged and revoked at handover. Production data access is minimised by design: most development runs against synthetic or anonymised data, with production access requested case by case.

Development environments are decommissioned and any copies of your data are deleted on a schedule agreed in the contract, with written confirmation. Anything you want retained for support is kept only under an active agreement.

Our practices are SOC2-aligned — access reviews, change management, encryption, logging, vendor management — and we build client systems to pass SOC 2 audits. Where your procurement requires a certified provider for hosting, we deploy into your certified environment rather than ours.

We treat client personal data as processed on your instruction, with purpose limitation, consent handling where you are the data fiduciary, and support for data-principal rights such as access and erasure. Where an AI system stores derived data about individuals, we design retention and deletion paths for it explicitly.

Yes — open-weight models deployed in your VPC or on-premise, with no data leaving your environment. Standard for banking, healthcare, and government clients. The trade-off is capability: the best open models are strong but not equal to frontier hosted models on the hardest reasoning tasks.

Technology & models

Plain explanations of the technology, and how we choose between options.

We are model-agnostic: Claude, GPT-4 and Gemini for frontier tasks, plus open-source models (Llama, Mistral) where cost or data residency favours them. A routing layer picks the cheapest model that passes your quality bar per task.

Retrieval-Augmented Generation: before the AI answers, the system fetches relevant passages from your approved documents and instructs the model to answer only from them. It is how we make assistants cite sources instead of guessing.

A chatbot answers questions in conversation. Automation executes a fixed process on structured triggers. An agent pursues a goal: it plans steps, calls tools, checks its own output, and handles cases that were never explicitly programmed. Agents cost the most to run and are worth it only when context spans systems and judgement varies.

Only when justified. Retrieval and prompt engineering solve most enterprise cases at lower cost and risk; we fine-tune when a task needs consistent style or domain behaviour that prompting cannot reach, or when a small fine-tuned model can replace an expensive frontier one at volume.

Evaluation suites: golden datasets scored automatically on every change, plus business KPIs tracked post-launch. Thresholds are agreed with the business before the build. No system goes live on impressions.

A set of real cases — typically several hundred — with agreed correct outcomes, plus a scoring function. Every change to a prompt, model or retrieval configuration runs against it automatically, so regressions are caught before release rather than by a customer.

Sending each step to the cheapest model that clears a measured quality bar for that specific step, rather than sending everything to a frontier model. In typical pipelines this removes 30–50% of model spend with no measurable quality change, because most steps are classification and formatting rather than reasoning.

Model providers can cache the stable portion of a prompt — system instructions, retrieved context, few-shot examples — so it does not have to be reprocessed on every call. It reduces both cost and latency significantly for applications that reuse context, which is most of them.

Three things together: ground the model in your content through retrieval so it is working from sources rather than memory; make it cite so claims can be verified; and instruct it to state clearly when the answer is not in the material. Then measure correct-refusal rate, which is the earliest signal that it has started guessing.

Usually a vector index, not necessarily a separate database. Postgres with pgvector is sufficient for most enterprise corpora and avoids a new system to operate. Dedicated vector databases earn their place at very large scale or with demanding filtering requirements.

For classification, extraction, summarisation and many retrieval-grounded tasks, yes — and they are dramatically cheaper at volume and deployable inside your environment. For the hardest multi-step reasoning they still trail frontier models. We benchmark both against your actual task rather than deciding on principle.

Context is everything the model can see when producing a response: your instructions, retrieved documents, conversation history and the current question. Every model has a maximum, and cost and latency rise with how much you use. Good retrieval matters precisely because it puts the right small amount into context rather than everything.

The evaluation suite makes upgrades routine: run the new model against the golden set, compare scores, and promote if it passes. Systems are built with the model as configuration rather than hard-coded, so switching is a change of setting plus an evaluation run.

A single internal service every AI request passes through, handling authentication, redaction, model routing, caching, rate limits and audit logging. Without one, each integration re-implements security independently and nobody can answer what the organisation is spending or sending.

Pricing & commercials

How we price, what things cost, and what happens if outcomes are missed.

Discovery is fixed-fee. Delivery is fixed-scope, fixed-price after discovery. Managed operations are monthly with SLAs. No open-ended time-and-materials surprises.

Focused deployments typically start around ₹25–40 lakh; larger agent systems and products are scoped individually. Every quote comes with an ROI model showing the assumptions behind the value, so you can challenge them.

KPIs are agreed before we build. If a launched system misses its committed metrics, we fix it at our cost — that guarantee is in our contracts. It is also why we decline work we do not believe will reach its target.

Yes — managed-operations retainers covering monitoring, improvement cycles, and new-capability sprints, with monthly reporting against your KPIs. Retainers are cancellable with notice and never a condition of the build.

Milestone-based: a portion at kickoff, the remainder tied to agreed delivery milestones. No large upfront lock-ins, and you can stop at any milestone boundary keeping everything produced to that point.

Discovery is fixed-fee and sized to scope — typically a few lakh for a focused single-workflow discovery, more for a multi-function enterprise assessment. It is priced so that it is worth running even if you then choose not to proceed.

Inference, hosting and any human review time. For a grounded chatbot, inference is a fraction of a rupee per conversation. For an operations agent, a few rupees per completed case. We model these at 10× and 100× your pilot volume before you commit.

No — model usage is billed by the provider directly to your account, so you see the real cost with no margin from us. We configure and optimise it, and we report on it, but we do not resell it.

Yes. Tell us the budget and we will tell you honestly what fits inside it, or that nothing worthwhile does. Scoping to a real constraint produces better decisions than discovering the constraint after a proposal.

Occasionally, for product builds where we believe strongly in the thesis, and always alongside a cash component rather than instead of one. It is the exception rather than an offer we lead with.

From your volumes and fully loaded costs, not headline salaries — with each assumption stated separately so you can change it and see the effect. Where we cannot substantiate a number we mark it as an estimate rather than presenting it as a calculation.

None from us. You own everything we build and pay only your own infrastructure and model-provider bills. Where a third-party tool is genuinely the right answer, we recommend it and you contract with them directly.

AI agents & automation

Questions specific to autonomous agents and document or process automation.

Three tests: does it require assembling context from several systems, does the judgement vary case by case, and does it end in an action? If any is missing, something simpler is cheaper and more reliable — a rules engine or a grounded chatbot will beat an agent on cost, latency and predictability.

As much as they earn, per decision category. Autonomy is granted only after a category records zero harmful actions across the golden evaluation set and the red-team suite, and each expansion has a documented rollback.

A configured threshold above which the agent cannot act alone — a refund value, an external communication, anything touching a regulated record. The action queues for a named human with the full reasoning trace attached. Thresholds are configuration, so they can be tightened the day something feels wrong.

Usually not. A single agent with a well-designed tool set is cheaper, faster and far easier to evaluate. We add a second agent only for a structural reason — an independent reviewer with no write permissions, or a security boundary between a high-permission actor and a low-permission researcher.

On typical business documents, 95–99% field-level accuracy, with confidence thresholds routing anything uncertain to a human queue so errors do not silently pass through. We prove the number on your own documents during a parallel run rather than quoting a benchmark.

No. We usually add an AI layer in front of existing bots so they receive clean, structured input. Your existing investment is preserved and value arrives in weeks instead of a rebuild.

Workflows above roughly 500 documents a month typically pay back within the first quarter. Complex documents or high manual handling cost lower that threshold. Below it, integration and change-management effort usually outweighs the saving, and we will tell you so.

They route to a review queue with uncertain fields highlighted against the source image, keyboard-first correction and one-click approval. Every correction becomes evaluation data. A poorly designed queue makes an excellent extractor unusable, so we design it as carefully as the extraction.

The automated pipeline processes live volume alongside your existing manual process for two to four weeks, with every disagreement adjudicated. It produces the real accuracy figure on your real documents — and reliably measures the manual baseline’s own error rate, which is typically 1–4% and has usually never been quantified.

Yes. Anything with an API, database, file drop or message queue can be a tool. Where no API exists we build a thin service in front of the legacy system rather than letting an agent drive a user interface, because screen automation is the most fragile integration available.

Monitoring flags it, the audit log shows exactly which step failed, and the case becomes a regression test. If it indicates a category-level problem, that category returns to supervised mode while it is fixed. This is designed before launch rather than improvised after.

A whitelisted action space with scoped credentials — no general shell, no arbitrary HTTP client, no write access to systems it was not commissioned for — plus per-task spend and time ceilings. That single constraint eliminates most catastrophic-failure scenarios.

Chatbots & voice AI

Conversational deployments: accuracy, channels, escalation and compliance.

Not one built on retrieval grounding with strict answer boundaries. If the answer is not in your approved content, it says so and escalates. Answers carry citations so any claim traces to a document, and we monitor correct-refusal rate as an early warning that it has started guessing.

On tuned retrieval over well-maintained content, 90–95% answer correctness on real customer questions with correct refusal above 95%. The ceiling is set by your documentation rather than the model, which is why we audit content before quoting a number.

Often yes, but remediation is real work and belongs before the build. Our knowledge readiness audit benchmarks retrieval against your actual content and reports what accuracy you would get today, what you would get after cleanup, and which documents are causing the damage.

Web, WhatsApp, Slack, Teams, Messenger and in-app — one brain, many channels. Formatting and available actions adapt per channel while knowledge, policy and evaluation stay shared, so every surface gives the same answer.

The escalation carries the full conversation, the customer record, what the assistant attempted and why it stopped, straight into your helpdesk. The agent picking it up starts from a summary rather than from scratch, so handling time on escalated contacts typically falls.

No — modern neural voices with natural prosody and interruption handling. Where callers do report an unnatural experience the cause is almost always timing rather than voice quality, which is why latency and endpointing get most of our tuning effort.

Under about 800 milliseconds from the caller finishing to the agent starting. Past roughly one second, callers assume they were misunderstood and start repeating themselves. Every architectural decision in a voice deployment is subordinate to that budget.

Yes: booking, rescheduling, payment links, identity verification, ticket creation, address changes — anything your APIs allow. Actions are latency-budgeted so they happen inside the conversation, and consequential ones are confirmed back to the caller before execution.

No. The agent sits alongside existing telephony via SIP, Twilio or your contact-centre platform, taking a defined slice of traffic. No number porting is needed to pilot and routing can be reverted instantly.

After-hours and overflow calls — those that currently reach voicemail. There is no comparison case where a human would have done better, so the pilot can only improve the caller experience. We expand into business-hours traffic once containment and QA scores are proven.

Yes, with disclosure. Several jurisdictions now require telling callers they are speaking to an AI system, and the EU AI Act requires it for interactions with people. We disclose on every call by default regardless of local requirement, and we capture recording consent verifiably.

We benchmark speech models against recordings from your actual caller base rather than trusting vendor claims, and tune noise robustness in the same pass. Where a particular accent or environment underperforms, those calls route to humans rather than degrading the experience.

AI products: apps & SaaS

Building AI-native web apps, mobile apps and SaaS products.

Streaming instead of spinners, sources instead of black boxes, and review-and-undo instead of irreversible actions. These patterns make probabilistic output feel controllable, and products that adopt them see markedly higher adoption of AI features.

Streaming end to end so users read while generation continues, model routing so simple steps use fast models, prompt caching, and optimistic UI on actions. We optimise time to first token at the 95th percentile, because averages hide exactly the sessions that lose users.

For most products, a hybrid of a seat price with a generous included allowance and transparent overage. Pure per-seat leaves margin exposed to heavy users; pure usage-based creates meter anxiety that suppresses adoption. The right answer depends on your usage distribution, which we model before pricing.

70–85% with disciplined architecture: model routing, caching, and pricing designed against the real usage distribution rather than the average. Products that skip this often run at negative margin on their heaviest cohort without knowing it.

Row-level security at the database so a query missing a tenant filter returns nothing rather than everything, retrieval indexes partitioned per tenant rather than shared with metadata filtering, tenant-keyed caches, and content-redacted logs.

SSO and SCIM, RBAC, exportable audit logs covering AI actions, a DPA, an accurate subprocessor list naming your model providers, architecture and data-flow diagrams, and a written answer that customer data is not used for training. Building these during the MVP avoids losing a quarter to them later.

Yes — quantised models run on-device via Core ML or TensorFlow Lite, working with no signal, at zero cost per inference and with data never leaving the handset. Capability is limited compared with cloud models, so most apps are hybrid: local for the fast, frequent, private path.

Yes, with preparation. Rejections are usually procedural: inaccurate privacy manifests or Data Safety declarations, missing content-moderation description, or age ratings that do not reflect generative capability. We prepare the compliance pack during the build, so submissions generally pass first time.

Yes — embedding a copilot or assistant into an existing React, React Native or native codebase is one of our most common engagements. Typically six to twelve weeks including evaluation, metering and security documentation.

Suggestion acceptance rate, edit distance on accepted output, regeneration rate, and abandonment during generation. Usage counts mislead — a user regenerating four times is engaged and unhappy.

Governance, compliance & regulation

AI policy, risk management, and the regulations that actually apply.

Quite possibly. It applies to organisations placing AI systems on the EU market or whose system output is used in the EU, regardless of where you are established. Transparency obligations catch more organisations than the high-risk rules do — chatbots and voice agents must disclose that they are AI.

Broadly, systems used in decisions that materially affect people: employment and recruitment, credit and insurance, education access, essential services, law enforcement and certain safety components. These carry obligations around risk management, data governance, logging, human oversight and technical documentation.

Tier it by risk. Low-risk internal uses operate under a clear usage policy with no approval needed; medium-risk uses get a lightweight documented review in days; high-risk uses affecting decisions about people get formal review with documented oversight and bias assessment. Heavyweight process on everything teaches people to route around it.

A documented understanding of what each model is used for, how its quality is measured, what happens when it degrades, who owns it, and what the fallback is. For regulated sectors it extends to bias testing, challenger models and periodic revalidation.

Yes, in an increasing number of jurisdictions, and we recommend it universally. Disclosure costs nothing and removes the entire category of "we were misled" complaints. It is required under the EU AI Act for interactions with people and under several US state laws for voice.

For systems affecting people we test outcomes across relevant groups on a held-out dataset before launch and monitor them after. Where disparities appear, we address them in the system design and document what was tested and found — which is both good practice and what regulators ask for.

A short document telling employees what they may and may not do with AI tools, what data classes may be used with which tools, and where to go for approval. Yes, you need one — and it needs to be paired with a sanctioned tool, because a policy without an alternative just moves usage to personal devices.

Provide a sanctioned path that is easier than the unsanctioned one. Prohibition alone reduces visible usage rather than actual usage. We size the current exposure during an audit — the number is usually higher than leadership expects — and it is generally the most persuasive case for funding the governed alternative.

Versioned process and prompt definitions, complete run records showing what the system did and when, documentation of human oversight points, evidence of pre-deployment evaluation, and an incident log. Systems built on orchestrated workflows generate this as a by-product rather than requiring it to be assembled.

Your organisation, which is exactly why we build named ownership, approval gates and audit trails into every system. Accountability cannot be delegated to a model, and any design that implies otherwise is a design problem rather than a legal question.

Getting started

Practical first steps, whether you are ready to build or still deciding.

With a discovery that scores opportunities on value and feasibility across your functions, inspecting real data rather than asking whether data exists. Four to six weeks, fixed fee, and the output includes what not to do — which clients frequently say saved them more than the recommendations.

A problem, not a solution. The workflows that consume the most time, the backlogs that never clear, or the thing your competitors do faster. We will tell you which are good AI candidates and which are better fixed another way.

Pick one workflow, measure its current volume, handling time and fully loaded cost, and model the value of a realistic improvement. A specific, measurable case for one workflow gets funded far more often than a thematic AI programme, and it creates the reference for everything after.

A named production owner from day one — someone whose job depends on the system working next year and who has veto power over its design. In our client data, projects with one reach production roughly four times as often. It is the strongest signal we track and it costs nothing.

Both, usually in that order of eventual importance. Partner to get the first systems live and to establish the patterns; hire to own and extend them. We build for handover from the first commit precisely because a permanent dependency on us is a bad outcome for you.

Keep the model as configuration rather than architecture, invest in evaluation suites rather than prompt cleverness, and build the durable parts — data pipelines, integrations, guardrails, orchestration — properly. Models change every few months; those foundations do not.

One bounded workflow with a measurable baseline, a named owner and available data. A chatbot on one channel, one document type automated, or one exception category handled by an agent. First projects should be chosen for certainty of completion rather than ambition.

A feasibility spike produces a working prototype against your real data in two to three weeks. A first production deployment ships in four to eight. We show working software weekly from the first week rather than status decks.

Still have a question?

Ask an architect directly — we answer within one business day.

Contact us