Buying software

Adding AI to Business Software: Where It Helps

Most requests to add AI start from a demo. The useful question is narrower: which step done by hand today is reading, sorting, finding or drafting - and what it costs when it is wrong.

8 min read Plexowave

We build AI features into business software, and a fair share of the conversations we have about them end with a recommendation not to. That is not modesty. A model is a good tool for a specific kind of work, and an expensive, unpredictable one for everything else. Knowing which is which before anything is built is most of the value.

The pattern that goes wrong is familiar. A feature is chosen because it demonstrates well, usually a chat box, and it is added where it is visible rather than where the manual work is. Nobody defines what a correct answer looks like, so nobody can say whether it is working. Six months later it is still there and nobody uses it.

Four jobs it does well

Almost every AI feature that earns its place in business software is one of these four, and each of them replaces a person reading text and typing what they found.

  1. Reading documentsSupplier invoices, purchase orders, bank statements and delivery challans read into structured fields. The model reads; ordinary code then checks what it read - that line totals add up, that a GSTIN has the right shape, that the supplier exists. Most extraction errors are caught by those checks before anyone sees them.
  2. Sorting what arrivesEnquiries, tickets and emails routed to the right queue, tagged by urgency, or flagged when they need a senior person. The cost of a wrong guess is a re-route, which is what makes this a good first task.
  3. Finding things in your own recordsQuestions answered from your documents, policies and past records, with the source cited next to the answer. Without the citation it is a confident paraphrase nobody can check; with it, it is a faster search.
  4. Drafting for a personReplies, quotation covers, product descriptions and summaries of long threads, prepared for someone to correct and approve. The saving is real when correcting a draft is faster than writing from nothing, and only then.

Where it must not decide

Anything that is a calculation or a rule belongs in code. Payroll, GST, stock valuation, interest, discounts and eligibility have one right answer that an auditor can retrace, and a model produces a plausible answer that is usually right. Usually is not a standard anyone can file a return on.

  • Tax and statutory computations, even where a model could explain the rule perfectly well.
  • Approvals with a financial or legal consequence, such as credit limits, refunds and write-offs.
  • Anything where the same input must always give the same output.
  • Facts that are not in the data it was given. Asked about something it cannot see, a model fills the gap with something that sounds right.

Measure before you ship

The single habit that separates AI features that last from ones that get switched off is an evaluation set, built before the feature rather than after the complaints.

  1. Collect real casesFrom your own documents and messages, including the awkward ones - poor scans, handwriting, mixed languages, the supplier whose invoice layout changes every quarter.
  2. Write down the right answerFor each case, what a careful person would have produced. This is tedious, and it is the step that makes everything after it possible.
  3. Measure per field, not overallAn invoice reader that gets totals right and dates wrong has a specific, fixable problem. A single accuracy figure hides it.
  4. Set a thresholdBelow an agreed confidence, the case goes to a person instead of through. The threshold is a business decision, and it should be yours.
  5. Re-run on every changeA new prompt or a new model becomes a measured comparison against the same set, rather than a change nobody dares to make.

The human gate

Every AI feature needs a decision about what happens without a person. For most business tasks the right starting point is to draft for approval and never act unattended, and to widen that only as the measured numbers justify it.

The gate only saves time if checking is faster than doing. That means showing the source next to the output - the invoice image beside the extracted fields, the cited paragraph beside the answer - so a person confirms rather than redoes. A review screen that makes someone hunt for the original has quietly cancelled the saving.

Cost, speed and data

  • Pricing is per call, so the bill scales with volume. Estimate monthly volume and the size of each request before building, not after the first invoice.
  • Cache work that repeats, and route easy cases to smaller, cheaper models. Most volume is easy cases.
  • Set a latency budget. A feature that takes many seconds inside a busy screen will not be used, however good its answers.
  • Decide what data may leave your infrastructure, and redact what need not be sent. Personal data in a prompt is still personal data under India's DPDP Act.
  • Check the provider's terms on retention and training. Enterprise API terms generally exclude training on submitted data; where that is not enough, an open-weight model can run on your own servers.

A sensible first project

Pick one task that is frequent, low-risk and easy to check. Reading supplier invoices into a purchase entry that a person approves is a good example: the volume is real, a mistake is caught at approval, and the time saved is easy to measure. Build the evaluation set, ship it behind the gate, and look at the numbers after a month.

If they are good, the next task is easier to justify and quicker to build, because the evaluation habit and the review screen already exist. If they are not, you have learned it cheaply - and sometimes what you learn is that the problem was a badly designed form, which needed no model at all.

Questions

Common questions.

Which AI model should we use?

The one that measures best on your evaluation set at a cost you can carry. Models change quickly, so the integration should make swapping one a configuration change, and the evaluation set turns that swap into a comparison rather than a guess.

Can AI read Hindi, Gujarati or handwritten documents?

Printed text in Indian languages often reads well with current models. Handwriting and poor scans vary a great deal from one document to the next. Your own samples in the evaluation set are the only reliable answer for your documents.

Will it replace the people doing this work now?

In practice it moves them from typing to checking, and raises how much one person can get through. The judgement stays with a person, which is where it should stay for anything with a consequence.

How long does a first AI feature take?

Typically seven to ten weeks from finding the work to a measured pilot: a week to find the right task, a week to build the evaluation set, three to five to build and measure, and two to three on live volume behind the gate.

Working through this decision?

Describe the problem rather than the solution. If the answer is a product you can buy, or no software at all, we will say so.

Start a project