Vendor selection starts with finding candidates and quickly moves to a harder question: how will each vendor perform when an urgent release lands, context is missing, or your internal reviewers disagree about what “good” looks like?
A polished sample, an ISO certificate, a low word rate, and an impressive demo are useful signals. Evaluate them alongside the people, workflow, technology, full costs, and safeguards behind the service. Ask for evidence that the promises will hold up in day-to-day work.
How to choose a localization vendor in eight steps
To choose a localization vendor:
- Find the source of the problem: the language service provider, localization platform, internal workflow, source content, or governance.
- Describe how localization should work across your content, locales, release cadence, quality standards, systems, data risks, and internal responsibilities.
- Choose the right sourcing model: one global provider, regional specialists, an integrator, a hybrid model, an in-house team, or a platform-only setup.
- Choose the lightest request format that will make vendors comparable: a clear brief or RFQ for defined work, and an RFI or RFP for a broader operating solution.
- Compare proposals by total cost of ownership, including every cost around the word rate.
- When the stakes justify it, run a controlled pilot using representative content, the proposed production team, and the proposed production workflow.
- Remove candidates that fail a non-negotiable requirement before scoring the rest. Put the evidence behind the winning offer into the contract.
- Plan transition, asset portability, and eventual exit before signing.
Use the same sequence whether you are replacing an incumbent or adding a specialist. Bring in product, procurement, security, and other stakeholders where their input affects the decision.
A lighter path for mid-sized teams
The framework scales down to a lean selection process. Match the effort to the risk, complexity, and cost of reversing the decision. A mid-sized company handling regulated content or a complex migration may need the full version. A larger company buying a straightforward, low-risk service may need far less.
If your scope is clear and the risk is moderate, keep the logic and trim the paperwork. One workable version starts with three candidates, takes two into finalist conversations, and scores them against the five criteria that matter most. Adjust those numbers to fit your decision.
- Write a one-page brief covering the content, languages, expected volume, turnaround, tools, quality expectations, and any restrictions on confidential material, personal data, or AI.
- Screen out vendors that miss a non-negotiable requirement. Ask three credible candidates to price the same realistic scenario and explain how they would run it.
- Take the strongest two into final conversations. Meet the people who would manage the work, and ask what happens when feedback conflicts, a deadline moves, or a team member is unavailable.
- Compare the full cost, including the time your team will spend preparing files, answering questions, reviewing work, and managing the relationship.
- If the choice is close, or a poor decision would be costly to unwind, run the same focused, representative pilot with both finalists.
- Score the finalists on the five criteria that matter most to you. Handle non-negotiables as pass/fail gates. Put the essentials in writing: scope, pricing assumptions, turnaround, quality and escalation, approved AI use, asset ownership, and a workable exit.
A lean process may take several weeks once the brief is ready. Allow time for comparable responses, finalist conversations, and, if needed, a focused pilot. A process that ends after the finalist conversations can move faster. Contracting, security or legal review, integrations, regulated content, hard-to-source languages, and migration can extend the schedule.
For a small, easy-to-reverse engagement, comparable quotes, a reference check, and a limited initial engagement may be enough. Add the full RFP, scoring, pilot, and transition work when a poor choice would be hard or expensive to undo.
First, decide whether you need an LSP, a TMS, or both
An LSP and a TMS solve different problems:
| Term | What it is | What it usually manages or handles |
|---|---|---|
| Language service provider (LSP) | A company that supplies and manages linguistic services | Linguists, editing, quality control, project management, capacity, and delivery |
| Translation management system (TMS) | Software for orchestrating localization work | Connectors, workflow automation, translation-memory storage, terminology, permissions, and reporting |
| Managed localization solution | A service arrangement that combines people, workflow, and technology | Some or all of the above, with one party accountable for the overall process |
Good translations paired with manual file handling, poor status visibility, and copy-and-paste delays point to the platform or integration. Smooth automation paired with drifting terminology and voice points to the provider and reviewer model. Unclear priorities and unresolved feedback point to governance. A new tool or vendor will leave that bottleneck in place.
Before launching a search, pin down the source of the problem. Alconost's localization-platform comparison covers software selection. This article covers the service partner and the way the work will run.
When should you review or replace a localization vendor?
It may be time for a formal review when several of these problems keep coming back:
- Localization repeatedly delays releases or campaigns.
- The same errors recur after feedback has been accepted.
- Results depend on one project manager or linguist, with no credible backup.
- Internal reviewers spend too much time rewriting, reconciling terminology, or chasing status.
- Your content mix, locale footprint, volumes, or risk level has changed materially.
- The vendor cannot support necessary repositories, design tools, content systems, or automated workflows.
- Reporting does not let you explain cost, quality, throughput, or rework to stakeholders.
- Security, privacy, or AI-use questions receive vague answers.
- Translation memories, glossaries, reports, or workflow configurations are hard to export and reuse.
- Commercial terms have become difficult to compare with the value delivered.
These problems give you a reason to investigate. Start with a baseline: missed deadlines by cause, corrections by category, internal review hours, engineering defects, turnaround by content type, and total spend. That baseline helps trace each failure to the provider, source content, late product changes, unavailable reviewers, or unrealistic planning. It also gives the next provider something concrete to improve.
Step 1: Define how localization should work
The question “Can you translate these languages?” produces a broad sales answer. Give candidates a picture of how localization needs to work inside your business.
Bring the right stakeholders in early
The selection team should represent the work the vendor will actually touch:
| Stakeholder | Questions they should own |
|---|---|
| Localization or content operations | Workflow, linguistic quality, terminology, reporting, escalation, and governance |
| Product, engineering, or design | Formats, repositories, builds, screenshots, character limits, internationalization, and release cadence |
| In-market reviewers or business owners | Local suitability, regulated terminology, brand voice, and acceptance |
| Procurement and finance | Proposal comparability, total cost, commercial risk, and contract structure |
| Security, privacy, or legal, when relevant | Data sensitivity, access, retention, AI use, contractual controls, and incident response |
You rarely need every function represented in every meeting. In a smaller company, one owner can gather input as needed. Cover the relevant questions before the decision and involve specialists where their expertise matters.
Early input from security and engineering reduces rework at the finalist stage. The vendor also needs to meet the technical and security requirements for onboarding.
Build a requirements baseline
Give candidates the same realistic description of:
- Content: products, file formats, repositories, content types, and source-language quality.
- Locales: current and planned languages, including regional variants and lower-volume markets.
- Demand: historical volume, seasonality, release peaks, urgent work, and expected growth. One annual word count is not enough.
- Risk: which content is customer-facing, regulated, safety-relevant, legally sensitive, or low consequence.
- Quality: the business definition of acceptable output, recurring error patterns, review method, and escalation rules.
- Workflow: intake, preparation, translation, editing, automated checks, in-context review, sign-off, delivery, and updates.
- Technology: TMS, repositories, CMS, design tools, issue trackers, APIs, and identity or access requirements.
- Data: confidential content, personal data, pre-release material, geographic restrictions, and permitted uses of machine translation or generative AI.
- Ownership: what the vendor owns, what internal teams own, and how feedback becomes reusable terminology or translation-memory updates.
- Success: what must be measurably better in six to twelve months.
If you have no documented workflow, sketch one. Mark every manual handoff, waiting point, reviewer, and system boundary. Alconost's overview of the localization process can help you spot stages that are easy to miss.
Set the level of care by content risk
One practical approach is to group content into tiers:
- High consequence: legal, safety, financial, medical, security, or reputation-sensitive content.
- Brand and revenue: product UI, onboarding, website pages, campaigns, and sales materials.
- Operational: help content, release notes, internal enablement, and knowledge-base material.
- High-volume or short-lived: support conversations, user-generated content, or rapidly expiring material.
Let those tiers determine human involvement, reviewer qualifications, turnaround, automated checks, acceptance rules, and whether MT or AI is allowed. This helps control spending on low-risk material and gives higher-risk content the safeguards it needs.
Step 2: Choose the appropriate sourcing model
Choose the model around domain risk, locale coverage, and how much coordination your team can realistically take on. Each option has its own trade-offs.
| Model | Often fits when | Main trade-off to test |
|---|---|---|
| Primarily in-house | Localization is strategically sensitive and demand is stable enough to support dedicated roles | Hiring, coverage, peaks, specialist depth, and tooling remain your responsibility |
| One multi-language LSP | Central accountability and consistent governance matter across many locales | Verify local-market depth, team continuity, and coverage gaps across locales |
| Regional or domain specialists | A few markets or specialist content types require deep expertise | Your team must coordinate more vendors, assets, reporting, and handoffs |
| Managed integrator | You want one layer to coordinate several providers, platforms, or workflows | Confirm transparency, markup, decision rights, and direct access to data and assets |
| Hybrid or multi-vendor | You need redundancy, benchmarks, or different workflows by content tier | Governance and fair performance comparison become more complex |
| TMS plus internally managed resources | You have strong localization operations and mainly need orchestration technology | Your team retains supplier management, linguistic accountability, and internal ownership |
Company size alone tells you very little. Give candidates one or two realistic scenarios, such as a synchronized launch across 18 locales or an urgent update to regulated content. Ask them to walk you through who does what, where the work happens, and how problems are escalated.
Our overview of roles in a localization team can help you see which responsibilities should stay in-house and which a partner needs to cover.
Step 3: Choose an RFI, RFQ, or RFP and ask for evidence
The terms are sometimes used loosely. The US General Services Administration provides a useful distinction: an RFI gathers market information, an RFQ requests pricing for defined work, and an RFP asks vendors to propose a complete solution. Your procurement team may use slightly different language.
| Document | Use it when | Expected output |
|---|---|---|
| RFI | You are still learning which operating models and capabilities are feasible | A market map and a defensible shortlist |
| RFQ | The requirement is standardized enough for like-for-like pricing | A comparable quotation against the same specification |
| RFP | Workflow, quality, technology, security, and transition matter alongside price | A proposed solution, evidence, implementation plan, and commercial model |
For a complex program, a short RFI can screen the non-negotiables before several vendors spend days preparing full proposals.
Make the request comparable
Give every bidder the same demand scenario, content sample, file formats, service scope, pricing sheet, and evaluation schedule. Label mandatory requirements and preferences clearly. Give exceptions and assumptions their own section in the response template.
Ask vendors to separate:
- included and optional services;
- one-time and recurring costs;
- pass-through costs and markups;
- human translation, editing, review, MT post-editing, and AI-assisted workflows;
- project management, engineering, testing, technology, onboarding, and migration;
- standard and expedited turnaround;
- volume assumptions, minimum fees, and discount bands.
Use the claim, evidence, contract test
Every claim that affects the decision needs proof. It also needs a home in the agreement if you select that vendor.
| Vendor claim | Evidence to request | If selected, capture in… |
|---|---|---|
| “We provide consistent quality” | Defined process, anonymized quality reports, reviewer qualifications, corrective-action example, and pilot results if a pilot is used | Acceptance method, threshold, reporting, remediation, and retest provisions |
| “We can scale for launch peaks” | Capacity scenario, named roles, backup plan, and a comparable client reference | Capacity assumptions, lead times, service levels, and escalation |
| “Our integration is seamless” | Demonstration using your format or sandbox, implementation plan, support model, and known limitations | Responsibilities, milestones, support levels, change control, and acceptance |
| “Your content is secure” | A written route and controls proportionate to the content. For personal or higher-risk data, this may extend to a DPA, subprocessor details, assurance reports, retention, deletion, and incident procedures | Approved route, controls, responsibilities, change notice, and deletion |
| “We use AI responsibly” | The proposed models and providers, purpose, eligible content, retention, training/evaluation terms, human review, and change controls. If AI is prohibited, request a documented human-only route. | Allowed and prohibited uses by content tier, approved providers, notice, and evidence |
| “Your assets remain portable” | Sample exports and an export rehearsal covering memories, terms, metadata, and reports | Ownership, formats, frequency, assistance, timing, fees, and end-of-contract obligations |
If a vendor cannot demonstrate an important promise or put it in the agreement, leave it out of the score.
Step 4: Evaluate capability across the whole production system
Keep the criteria consistent across the written proposal, finalist interviews, reference calls, and any pilot you run. Each stage will test the same claims.
Linguistic quality and domain fit
Look at how the vendor builds and supports the language team for each content tier, from recruitment and testing to briefing, assignment, and continuity. Useful evidence includes:
- subject-matter and locale qualifications relevant to your risk;
- documented use of style guides, termbases, translation memories, and reference material;
- editing and review responsibilities that are unambiguous;
- a defined error taxonomy, severity model, sampling plan, and process for resolving reviewer disagreements;
- evidence that feedback leads to lasting process improvements across future work.
Standards can support due diligence. ISO 17100 specifies requirements for translation services, ISO 18587 covers full human post-editing of machine-translation output, and ISO 5060 provides guidance for evaluating translation output. Check which legal entity, locations, and services the certificate covers, then test the proposed team against your requirements. Alconost's guide to ISO-certified translation explains the scope and limits of certification.
For product localization, distinguish linguistic review from in-context testing. A string can be correct in isolation and still appear clipped, misplaced, unreadable, or wrong for the screen. Linguistic quality assurance and proofreading catch different problems.
Team, governance, and capacity
Ask who would actually run your account. A company-wide headcount tells you little. Identify the program owner, project managers, language leads, engineers, quality owner, security contact, and backups. Then ask:
- Which named people will work during onboarding and after launch?
- Which roles may be subcontracted?
- How are linguists added or replaced?
- What happens when a key person is unavailable?
- How will peaks, urgent work, and new locales be staffed?
- What meeting cadence, reports, and escalation channels are proposed?
- Who has authority to stop a release or approve an exception?
A credible capacity plan explains staffing, backup coverage, and constraints. A broad claim such as “We support every language 24/7” needs those details behind it.
Workflow, technology, and integration
Follow one job all the way from the source system to published localized content. A connector logo on a slide is only the beginning. Ask candidates to show:
- how jobs are created, routed, prioritized, and updated;
- how context, screenshots, character limits, variables, and do-not-translate rules reach linguists;
- how translation memories and termbases are maintained;
- how branching, continuous updates, and source changes are handled;
- which automated checks run and who resolves the findings;
- how status and cost data return to your systems;
- what happens when the connector or API fails;
- which implementation and ongoing support work belongs to each party.
Portability depends on the setup. Test exports, API access, formats, ownership terms, migration effort, and contractual exit support for any platform under consideration.
Confidentiality, privacy, and the AI data route
Scale the review to the content. Public marketing copy with no personal data may need an answer of only a few lines. Confidential material, unreleased source code, customer-support logs, and regulated content call for a map of the full data route: every system, provider, location, and person the content passes through.
Broad claims about AI use or model training need more detail. Match the depth of the questions to the content risk. Confirm:
- Which people at the LSP or its subcontractors, and which platforms or model providers, can access the content?
- For what purpose does each party process it?
- In which locations is it stored and processed?
- How long are source content, outputs, prompts, logs, and backups retained?
- Can any of it be used for model training, evaluation, or service improvement?
- What access controls apply, and what evidence supports them?
- Can workflows differ by content-risk tier?
- How will the buyer be notified before a provider, model, or data route changes?
- How is data exported, returned, or deleted at the end of the relationship?
Match the evidence to the risk. Depending on the project, this may include an NDA, secure transfer, named-team access, a customer-controlled platform, a human-only workflow, or private model deployment, together with clear retention and deletion terms. Personal data or stricter assurance requirements may call for a DPA, current subprocessor details, data-flow documentation, access-control records, incident and continuity procedures, and the scope of any ISO/IEC 27001 certificate or SOC 2 report.
Requirements for ISO/IEC 27001 or SOC 2 evidence vary by law, policy, customer commitments, insurance, and risk assessment. Check the controls and workflow proposed for your content whether or not a formal report is required. For higher-risk supplier reviews, the NIST supplier due-diligence guidance and NIST Generative AI Profile are useful starting points.
Alconost, for example, describes NDA-protected, human-only workflows for confidential game content, the option to work in a customer's localization platform, and private model deployment for sensitive use cases. These pages show which options may be available. Ask Alconost the same question you would ask any candidate: which people, systems, providers, retention terms, and contractual controls will apply to this content?
Have security and legal set the evidence requirements. Give them the actual project route and the controls proposed for it.
Resilience and references
On reference calls, go beyond general satisfaction and ask:
- What did the vendor underestimate during onboarding?
- How much internal work does the program require?
- How consistent is the production team?
- How does the vendor handle a repeated quality issue or missed deadline?
- What happens during volume peaks or staff changes?
- How easy is it to obtain complete assets, reports, and support?
Choose references with a similar content mix, locale footprint, risk profile, workflow, and scale. A famous logo tells you little if the underlying program looks nothing like yours.
Step 5: Compare the full cost of each proposal
Per-word pricing gets plenty of attention because it is easy to compare. A lower rate can sit beside higher project-management fees, platform subscriptions, engineering work, internal review, rework, and migration costs.
The Chartered Institute of Procurement & Supply defines total cost of ownership as the end-to-end cost across acquisition, use, and end of life. Applied to localization, a practical year-one formula looks like this:
Year-one TCO = production + management + technology + engineering + quality + onboarding/migration + buyer labor + risk/rework allowance + exit or end-of-term cost
Price every proposal against the same scenario
Use your actual demand where possible. If the records are incomplete, give every bidder the same clearly stated planning scenario. Include:
- new words, fuzzy matches, exact matches, and repetitions;
- per-word, per-character, hourly, project, retainer, and subscription fees;
- minimum charges and rush premiums;
- translation, editing, review, post-editing, transcreation, and testing;
- project management and vendor management;
- TMS, connector, seat, storage, API, and support fees;
- file preparation, engineering, build support, and screenshots;
- onboarding, migration, terminology cleanup, and training;
- internal requester, reviewer, engineering, procurement, and security time;
- expected rework and incident cost;
- asset export, transition support, and termination fees.
Alconost's guide to localization cost covers common production costs. Add the surrounding operational and transition costs to your procurement model.
A fictional example: the lower word rate loses on TCO
This illustrative calculation uses fictional figures created solely for the comparison.
Assume 100,000 source words are translated into 30 languages over a year, creating 3 million billable word units across all target languages. Of those units, 55% are new words, 25% are fuzzy matches, and 20% are exact matches or repetitions. Assume equivalent output quality from both vendors.
| Assumption | Vendor A | Vendor B |
|---|---|---|
| New-word rate | $0.110 | $0.085 |
| Fuzzy rate as share of new-word rate | 70% | 80% |
| Exact/repetition rate as share of new-word rate | 20% | 30% |
Vendor B's headline new-word rate is about 23% lower. Once both quotes are applied to the same demand and the operating costs are added, the picture changes:
| Fictional year-one cost | Vendor A | Vendor B |
|---|---|---|
| Linguistic production | $252,450 | $206,550 |
| Project management | Included | $24,786 |
| Platform | Included | $24,000 |
| Engineering | $12,000 | $20,000 |
| Quality and in-context review | $18,000 | $22,000 |
| Onboarding and migration | $8,000 | $15,000 |
| Buyer labor | $24,000 | $39,000 |
| Risk/rework allowance used in this simplified example | $0 | $0 |
| End-of-term export and transition | $4,000 | $8,000 |
| Illustrative year-one TCO | $318,450 | $359,336 |
In this fictional scenario, Vendor B costs about 13% more in year one despite its lower headline rate. Every offer needs the same assumptions, and internal labor still counts when it sits in another budget. The example excludes taxes, currency movement, annual price increases, and exceptional rework. Add any of those that would be material in your program.
The example uses a large volume to make the cost differences easier to see. TCO analysis also works at much smaller volumes. A simple sheet covering production, minimums, project management, platform fees, internal review time, onboarding, and exit may be enough.
Run the model a few ways: expected volume, a launch peak, slower growth, and a shift toward more reused or machine-assisted content. If one uncertain assumption flips the result, validate it before you choose.
Step 6: Run a production-representative pilot
A test translation shows the quality a vendor can produce in a small, controlled sample. A pilot goes further by putting the proposed team and workflow through a realistic assignment.
Size the pilot around the risks you need it to reveal. A short UI sample may uncover terminology and character-limit problems. Long-form brand voice, release engineering, and peak capacity require different tasks. Choose the content and pass criteria accordingly.
Design the pilot around a decision
Write down the question the pilot must answer. For example:
- Can the proposed team meet a weekly release cadence without creating engineering rework?
- Can the workflow protect confidential pre-release content while using approved automation?
- Can the vendor improve consistency across a mature translation memory with known defects?
- Can it handle regulated terminology and resolve reviewer disagreements?
- Can it launch a new locale and leave the assets ready for ongoing reuse?
Then build the pilot to answer it:
- Choose representative content and locales. Include realistic difficulty, recurring pain points, context dependencies, and the locales that matter to the decision.
- Give every finalist the same inputs. Use the same brief, style guide, glossary, memory, screenshots, source files, deadlines, and clarification rules.
- Test the proposed team and workflow. Include file preparation, handoffs, TMS behavior, automated checks, reporting, and delivery. Record which linguists, reviewers, project managers, and engineers took part.
- Document the data route. Record who can access the pilot content, which systems and providers process it, where it is stored, and how long it is retained.
- Use qualified, calibrated reviewers. Choose bilingual reviewers with relevant locale and domain expertise. Agree in advance on severity levels and hide vendor names where practical to reduce bias.
- Evaluate quality and operations. Score the output, questions, adherence to instructions, turnaround, engineering defects, reporting, escalation, and time required from your team.
- Resolve disagreements and test remediation. Discuss material scoring differences, record the decision, and give finalists a controlled opportunity to explain what caused a problem, propose corrective action, and show how they would prevent a repeat.
Paying for a pilot can support a realistic setup, especially when engineering or specialist work is involved. Its validity still depends on representative conditions, comparable inputs, transparent methods, and a clear decision rule.
For a closer look at the controls surrounding linguistic delivery, see Alconost's description of its translation quality-assurance process.
Use a quality model that matches the content
The Multidimensional Quality Metrics framework provides a configurable way to classify errors. Its scoring models show how categories and severity feed into a score. Choose the categories your content needs, define severity with concrete examples, and keep critical-error rules separate from the overall threshold.
Before anyone scores the work:
- test the model on content your organization already considers acceptable and unacceptable;
- train and calibrate evaluators;
- decide how reviewer disagreements will be resolved;
- separate errors caused by source content, instructions, tooling, and translation;
- define whether internal preference changes count as vendor errors;
- record the time your own team spends reviewing and resolving the work.
Set the passing score according to content risk, user impact, and what the pilot needs to prove. MQM and ISO standards leave that threshold to the organization using them.
Convert the pilot into operating commitments
Pilot results are relevant to production only when the conditions that produced them carry over. Put these points in the statement of work:
- the assumed team structure and substitution rules;
- approved workflow, systems, AI or MT uses, and data route;
- quality taxonomy, sampling, acceptance, reporting, and remediation;
- turnaround and capacity assumptions;
- escalation contacts and response expectations;
- implementation milestones and acceptance criteria;
- any limitations or dependencies discovered during the pilot.
Step 7: Apply gates before weights
Before scoring vendors, apply pass/fail gates to every non-negotiable requirement. This prevents a polished presentation or low price from mathematically offsetting an unacceptable failure in security, coverage, or portability. Reserve gates for requirements the business truly cannot compromise on.
Set pass/fail gates
Depending on the program, gates might cover:
- required locale, domain, or regulated-content coverage;
- legal, privacy, security, and approved data-route requirements that apply to the content;
- required system integration or file-format handling;
- ownership rights and usable export of the translation assets the buyer needs to continue the work;
- minimum pilot outcome, including critical-error rules, if a pilot is used;
- named accountable roles and an acceptable continuity plan;
- feasible onboarding, transition, and exit commitments.
Write each gate as an observable requirement. For confidential pre-release content, for example: “Files remain in our TMS, access is limited to the NDA-bound project team, and no MT or AI is used.” Another project may need a DPA and approved subprocessors. Match the gate to the risk and workflow.
Weight the remaining differentiators
The weights below suit a relatively complex program where workflow and data handling matter alongside language quality. Adjust them to your own risk profile. A straightforward program may weight these categories very differently.
| Comparative criterion | Example weight |
|---|---|
| Linguistic quality and domain fit | 25% |
| Operating team, governance, and capacity | 15% |
| Workflow, technology, and integration | 15% |
| Confidentiality, data handling, and approved AI use | 15% |
| Commercial model and normalized TCO | 15% |
| Transition, ongoing governance, and exit | 10% |
| Resilience and relevant references | 5% |
Score data handling by how well the proposed controls fit your requirements. A documented human-only workflow in customer-controlled tools may be the strongest option for some content. Certificates provide supporting evidence when their scope applies to the work.
Define the 0-to-4 scale for each criterion. For any given criterion, a score of 3 might mean that the vendor meets the requirement, provides the requested evidence, and accepts the related commitment.
Have evaluators score independently before the consensus meeting and cite the evidence behind each material score. Then test how stable the result is:
- Does the winner change if five percentage points shift between two criteria, keeping the total at 100%?
- Does the winner depend on one uncertain cost or volume assumption?
- Are closely scored vendors meaningfully different, or effectively tied?
- Did a stakeholder reward attractive extras that do not solve the stated problem?
- Does the ranking still make sense after any pilot and reference findings are included?
Treat the scorecard as a decision aid. The final choice still requires judgment.
Step 8: Put the evidence into the contract and plan your exit
The proposal and any pilot should feed directly into the statement of work, service levels, data terms, and implementation plan. If a vendor won on a particular team, workflow, data route, response time, or quality method, write it into the agreement.
Document what matters for this engagement, including:
- service scope, content tiers, locales, volumes, and dependencies;
- roles, named or key personnel, substitutions, and subcontracting;
- pricing assumptions, included services, change controls, and invoice detail;
- turnaround, capacity, response, escalation, and reporting commitments;
- quality measurement, critical errors, acceptance, remediation, and repeated-failure handling;
- approved systems, MT or AI uses, subprocessors, retention, and notification of changes;
- security evidence, incident duties, continuity, and recovery expectations;
- ownership and permitted use of source content, translations, translation memories, termbases, style guides, prompts, configurations, and reports;
- implementation milestones, responsibilities, acceptance, and rollback;
- renewal, termination, transition assistance, export formats, fees, deletion, and evidence of completion.
Design the transition before the award
A provider change can disrupt the release cadence you are trying to improve. Ask finalists for a transition plan based on your actual assets and systems, then evaluate it as carefully as the proposal.
At minimum, decide how you will handle:
- Inventory: memories, termbases, style guides, in-flight jobs, locale configurations, automation, credentials, reports, open issues, and contractual rights.
- Transfer and cleanup: export formats, duplicate or low-quality entries, terminology conflicts, metadata, ownership gaps, and validation.
- Setup and rehearsal: roles, access, integrations, content-tier rules, any pilot findings, training, and a dry run.
- Parallel operation or controlled cutover: which content moves first, how conflicts are avoided, and how performance is monitored.
- Rollback and acceptance: objective launch checks, decision owners, defect handling, and a credible fallback.
- Legacy exit: final exports, open-work reconciliation, credential revocation, data deletion, confirmation, and support obligations.
Portability needs attention throughout the relationship. Schedule usable exports, keep ownership explicit, and occasionally confirm that another qualified team could open and reuse the material. Test every important technology dependency and give it a practical, contractual exit path.
Once the relationship is live, continue the same discipline through reviews, corrective actions, capacity planning, and change control. Our guide to localization vendor management covers those routines. Outsourcing localization looks at the broader client-provider relationship.
Vendor red flags that deserve follow-up
These signals call for a more specific question before you decide:
| Signal | What to clarify |
|---|---|
| A flawless free sample with no production detail | Who performed it, whether they will remain on the account, and which workflow and tools were used |
| A low rate that does not fit the requested pricing template | Missing services, match logic, minimums, management, technology, engineering, review, and buyer workload |
| Broad certification claims | Issuing body, legal entity, locations, services, current validity, and certificate scope |
| A vague “secure AI” label where sensitive content is involved | The proposed models and providers, or a confirmed human-only route, plus access, location, retention, data use, and change notice |
| A long connector list without a relevant demonstration | Setup work, limitations, ownership, support, failure handling, and real implementation effort |
| A generic “native linguists” answer | Domain qualifications, selection, briefing, review, continuity, and measured performance |
| Reluctance to put a material proposal or pilot promise into the contract | Whether the claimed capability will exist in production and what happens if it falls short |
| Vague export or deletion language | Assets, formats, metadata, timing, fees, assistance, verification, and backup retention |
| Only marquee references | A reference with a comparable content mix, risk profile, locales, scale, and workflow |
Localization vendor selection checklist
Before awarding the work, confirm that your team has:
- Diagnosed whether the LSP, TMS, source content, internal review, or governance is the actual constraint.
- Defined target outcomes and a measurable current baseline.
- Segmented content and data by risk.
- Disclosed realistic demand, formats, systems, dependencies, and buyer responsibilities.
- Chosen a sourcing model that matches the coordination capacity of the internal team.
- Separated mandatory gates from weighted preferences.
- Given every candidate the same pricing and operating scenario.
- Converted important claims into evidence requests and proposed contract terms.
- Defined risk-based data-handling requirements and mapped every route through people, platforms, MT or AI systems, and other third parties.
- Normalized total cost, including internal labor, onboarding, rework risk, and exit.
- Matched the pilot and review method to the stakes, or documented why a limited initial engagement is enough.
- Evaluated operational behavior as well as translated output.
- Checked references against a comparable program.
- Checked whether small changes to weights or cost assumptions change the winner.
- Captured the winning claims and any pilot assumptions in the agreement.
- Defined transition, rollback, asset ownership, recurring exports, deletion, and exit assistance.
Frequently asked questions
What is the difference between a localization vendor and a localization platform?
Do we need a localization RFP?
How should we test a translation vendor?
How long should a localization pilot be?
Is the lowest per-word rate usually the cheapest option?
Does ISO certification guarantee translation quality?
Does a localization vendor need ISO 27001 or SOC 2?
Who should own translation memories and glossaries?
Choose the vendor that works beyond the proposal
Evaluate the language list, word rate, features, and presentation in the context of the operating model. Choose a setup your team can run when deadlines tighten. Where the risk warrants it, verify that fit in a realistic pilot. Make sure you can leave without losing your content or language assets.
A strong proposal earns the vendor a closer look. Award the work once the evidence shows that its setup will work inside your organization.
If you are evaluating a partner for software, apps, games, websites, or marketing content, you can also explore Alconost's localization services and localization company model, or tell us about the workflow you are trying to build.































































































































































