AI consulting in Israel

Four different purchases are sold under one name. Which one you need first, what each actually produces, and how to tell an engagement that will change something from one that will produce a deck.

In short

Decide which of four things you are buying before comparing anyone: strategy, architecture, delivery, or governance. They have different deliverables and different right moments. An engagement that does not name which one it is defaults to strategy, because strategy is the cheapest to deliver and the hardest to be held to. And the single question that separates useful from decorative: what number was measured before the work started. Impact scores without a baseline are opinions arranged in a table.

Four purchases, one name

What The question it answers What it leaves behind Right moment
Strategy Which problems are worth attacking at all A short ordered list, each item with the measurement that would prove it worked Only when you cannot name the task you want removed. Otherwise skip it
Architecture How a system that touches your data should be built, and where the boundaries are Where data flows, what the model may see, what is deterministic, what is escalated Before the first build, and again before the first system goes near customers
Delivery Build it Working software, and the obligation to run it After the boundary above is written down. Before that, you are paying to discover it twice
Governance What the organisation is allowed to do, and how that is evidenced Rules with versions, an audit trail, and a named owner per decision type In parallel with architecture, not after. Retrofitting it changes where data flows

The last row is the one most often sequenced wrongly. Governance is treated as a compliance step at the end, but what it actually constrains is architecture: whether data leaves the organisation, whether a decision has to be explainable, whether a model version must be recorded. Decided late, those constraints arrive as a rebuild.

The question that sorts good engagements from bad

One question, and it is not about technology: what number will be measured before anything starts?

The common deliverable in this market is a list of opportunities scored by impact and effort. It looks rigorous and it is unfalsifiable. Nothing in it can be checked in six months, because nothing was measured at the beginning, so there is no baseline to compare against. Any outcome can be described as a success.

The alternative is unglamorous. Pick the task you want removed. Measure how long it takes today, how many people touch it, and how often it happens. Three numbers, collected in a week, without a consultant. Everything afterwards is judged against them.

And a second question worth asking. "What would make you tell us not to do this?" An engagement with no answer has only one possible conclusion, which means the conclusion was not produced by the work.

Conversational AI specifically

Several of the searches that reach this page name conversational AI, so it deserves a direct answer rather than a general one.

A vendor can deliver a working assistant. That is no longer the hard part. The hard part is the boundary, and it is four decisions:

  • What it answers from your own content, and what it must not answer at all. An assistant with no refusal behaviour will invent an answer rather than decline, and it will do so confidently.
  • What happens when it does not know. Escalation to a person is a feature that has to be designed, including what the person receives. "It says it cannot help" is not escalation.
  • How a wrong-but-confident answer is detected. Not by a user complaining. By sampling real conversations against sources, deliberately, on a schedule.
  • Who is accountable for what it tells a customer. This has a legal dimension before it has a technical one, and it should be settled in writing before launch rather than after the first incident.

If an assistant needs to answer from data inside your own systems rather than from your public pages, that is a different architecture with different access controls. Related: integrations.

What a governance deliverable actually contains

Governance is the row most often bought and least often specified, so it is worth saying what a useful one looks like. Four artefacts, and if an engagement produces a policy document instead of these, it produced reading material.

  • A list of decision types, each marked allowed, allowed with a human, or not allowed. Not principles. Named decisions: refuse an application, set a price, answer a customer about their account, draft an internal summary. The value is in the specificity, because that is what an engineer can implement and an auditor can check.
  • Rules with versions. If a policy changed, a decision from last year has to be explainable under the policy that applied then. Systems that stored only the current rules cannot justify their own history, and that is exactly what gets asked.
  • A data map in one page. What leaves the organisation, to which service, under which agreement, and what is retained there. One page, not an appendix, because its purpose is to be read by people who will not read an appendix.
  • A named owner per decision type. Not a committee. The person who answers when that decision goes wrong.

The fourth is the test of whether the rest is real. A policy with no named owner describes an intention; the same policy with a name against each line is a commitment someone has accepted.

And one question that reveals how an engagement will end. "Who maintains this after you leave?" Governance is not a document that stays true. Models change, vendors change terms, the business adds a use case. If no one inside owns the list, it is accurate for about a quarter.

What the Israeli context adds

Hebrew output needs a native signature. Not review by a speaker, sign-off by someone who writes the language professionally. An assistant producing slightly wrong Hebrew reads as careless in a way that is invisible to the team that built it, and it is customer-facing from day one.

Multilingual is the normal case, not an edge case. The same user may read Hebrew, work in English and receive content in Russian. Language belongs to the person and to each message, not to the deployment, and a missing translation must be visible as missing rather than quietly filled with another language.

Jurisdiction is a design input. Where data may sit, and whether a decision affecting a customer has to be explainable, are legal questions that determine architecture. A consultancy that treats them as a later checklist will have designed around assumptions it did not test.

How I work on this

Practice since 2004, work on AI systems since 2023, projects for clients in fourteen countries. I am an independent architect, which places me in the second and fourth rows of the table above and deliberately not in the third. I do not sell delivery capacity, so "this does not need AI, it needs a process fixed" and "your volume is too low to justify this" are conclusions that cost me nothing, and in measured cases they are the common ones.

What a review produces: the three baseline numbers for the task in question, the boundary written down as decisions rather than features, the governance constraints that actually bear on architecture, and an order of work with the risk named per step. What it does not produce is a scored opportunity list before anything has been measured.

Where to start

Name one task you want removed, and say how long it takes today. If you cannot answer the second part, that measurement is the first piece of work, and it does not need me.

Get in touch  |  Related: choosing a software partner in Israel, integrations, how I work

SLAtech LTD