Choose a custom web application development company by testing how well it can turn one real business workflow into a traceable delivery plan. Give every serious candidate the same scenario, then compare the users, records, permissions, states, integrations, failure paths, acceptance tests, and handoff evidence they identify. A polished portfolio shows presentation quality; this exercise shows whether the team can understand, build, verify, and operate your application.
This framework is for a business comparing partners for a portal, platform, internal system, booking workflow, marketplace, CRM or ERP function, or another custom web application. It does not replace technical, security, privacy, accessibility, legal, or procurement review for the specific project. It gives non-technical decision-makers a consistent way to ask for evidence before selecting a team.
Give every candidate the same operating scenario
Do not begin with “show us your technology stack” or “how many developers do you have?” Start with one important workflow the application must support. Describe:
- Trigger: what starts the work;
- User: who is acting and on behalf of which organization, team, or account;
- Record: which request, order, case, document, payment, project, or account is involved;
- Authority: where the correct data comes from and who may change it;
- Decision: what must be approved, rejected, calculated, or confirmed;
- Exception: what can prevent completion; and
- Evidence: what proves the workflow finished correctly.
At AP Works, we use this operating statement as a traceability test. A candidate should be able to follow the scenario from business event to interface, data rule, permission, integration, failure state, test, and production evidence. The purpose is not to get free architecture work. It is to see whether the partner asks the questions that make the eventual scope dependable.
If you are defining a portal rather than selecting a partner, the customer portal requirements framework goes deeper into organizations, records, actions, states, approvals, and audit evidence.
Compare a seven-part evidence chain
| Evidence | What a useful response contains | Weak signal |
|---|---|---|
| Workflow boundary | The first and last event, responsible users, systems involved, and what is explicitly outside the release. | A feature list with no start, finish, or accountable owner. |
| Records and authority | Important data objects, source-of-truth systems, ownership, validation, retention, and migration questions. | “We will connect the database” without defining which system is authoritative. |
| Permissions and states | Who may view or change each record, in which state, with approval and negative cases included. | Roles exist only as different navigation menus. |
| Integration behaviour | API availability, credentials, limits, duplicate handling, timeouts, retries, reconciliation, and a manual recovery route. | An integration logo treated as proof that the workflow will work. |
| Acceptance and QA | Testable outcomes for normal, failure, mobile, keyboard, accessibility, security, and recovery paths. | “QA included” with no journeys, environments, or acceptance owner. |
| Release and operations | Staging, deployment, rollback, monitoring, backups, support boundaries, incident ownership, and change process. | Work ends when production deployment succeeds once. |
| Ownership and handoff | Contract terms plus repository, environments, accounts, documentation, licences, exports, access recovery, and another team’s ability to operate the system. | “You own the code” without practical control of the running application. |
The chain matters more than a particular document format. A small project may use a concise workflow map and acceptance checklist. A regulated, multi-role, or integration-heavy application may need formal specifications and specialist review. The key is that important business requirements remain traceable through delivery.
Use a weighted partner scorecard
Score every candidate against the same evidence. Adjust the weights before proposals arrive so the most persuasive presentation does not silently change the decision criteria.
| Criterion | Weight | What earns the score |
|---|---|---|
| Workflow and domain understanding | 20 | Finds the real operating boundary, users, decisions, exceptions, and adoption constraints. |
| Delivery evidence and traceability | 20 | Connects requirements to designs, implementation, acceptance tests, and launch evidence. |
| Data, integration, and privacy reasoning | 15 | Identifies authority, access, minimization, migration, vendor dependencies, and failure recovery. |
| Security and accessibility approach | 15 | Turns applicable standards and risks into scoped requirements and complete-journey tests. |
| Ownership and operational continuity | 15 | Defines repositories, accounts, environments, documentation, support, recovery, and transfer. |
| Team and communication fit | 10 | Names the people doing the work, decision owners, review cadence, and escalation route. |
| Commercial clarity | 5 | Separates assumptions, inclusions, exclusions, optional work, third-party costs, and change control. |
A total score is a decision aid, not a substitute for judgement. A lower-risk internal tool and a public application handling sensitive information should not use identical thresholds. Record the reason for material deductions and any condition that must be resolved before contracting.
Ask who will actually do the work
The sales team, named expert, and delivery team may not be the same people. Ask who will lead product discovery, interface design, architecture, development, QA, deployment, and post-launch support. For each role, clarify availability, decision authority, subcontracting, and what happens if that person changes.
Relevant work is more useful than a large logo wall. Ask a candidate to walk through one comparable decision: the constraint, alternatives considered, evidence used, tradeoff selected, failure case tested, and what the client received at handoff. The example does not need to match your industry exactly, but the reasoning should transfer to your risk.
Make discovery produce reviewable deliverables
“Discovery” can mean a sales call, a paid project phase, or several weeks of product definition. The label is not enough. State what the phase will produce, who owns the outputs, which decisions it is expected to resolve, and whether another qualified team could use the material.
Depending on the project, useful discovery outputs may include a workflow map, user and organization model, record inventory, permission matrix, state diagram, integration register, data migration plan, annotated interface prototype, architecture decision record, risk register, release boundary, backlog, acceptance plan, and implementation estimate.
Also define the exit. A discovery phase should say whether the next decision is to build, revise the scope, buy an existing product, run a smaller pilot, or stop. If a spreadsheet or SaaS product still fits, the spreadsheet-to-application decision guide can help test whether custom software is justified before partner selection continues.
Require security claims to become requirements
“Secure,” “enterprise-grade,” and “best practice” are not acceptance criteria. Ask which risks and requirements apply to your application, where they are recorded, how they influence design and development, and what evidence will be delivered.
NIST’s Secure Software Development Framework is intentionally high level and can provide a common vocabulary between software producers and purchasers. The OWASP Application Security Verification Standard provides testable web-application security requirements and explicitly supports procurement and contract use. Neither standard turns a generic checklist into a security guarantee; the project still needs a risk-appropriate scope and qualified review.
Ask how the team handles authentication, record-level authorization, secrets, dependency review, environments, logging, backups, vulnerability reports, remediation, and incident responsibilities. For high-risk or regulated work, identify the required independent assessment before the contract is signed.
Test accessibility across the complete workflow
Accessibility cannot be proven by one automated score or a component library claim. Define the applicable standard, target level, representative users, assistive technologies, testing responsibility, defect process, and acceptance evidence.
WCAG 2.2 uses testable success criteria and treats a multi-page sequence as a complete process: if one step in the sequence does not conform at the claimed level, the process does not conform at that level. For procurement, that means testing the entire registration, approval, booking, payment, document-submission, or account-recovery journey—not only the dashboard landing page.
Separate legal ownership from operational control
A contract may grant rights to custom source code while the business still lacks practical control of the application. Before selection, map both layers with legal counsel where appropriate.
- Legal layer: custom code and design rights, pre-existing tools, open-source and commercial licences, data rights, confidentiality, and permitted reuse.
- Operational layer: repository administration, cloud and domain accounts, databases, deployment process, environment configuration, backups, logs, third-party services, billing, documentation, and recovery contacts.
Managed operation is not the same as lock-in. A partner can run infrastructure and support while the client retains agreed ownership, visibility, export rights, and a tested transition path. The proposal should state which responsibilities remain with the partner, which belong to the client, and what changes during a handoff.
Compare estimates through assumptions, not only totals
Two estimates rarely describe the same application. Ask each candidate to identify the release boundary, assumptions, exclusions, client responsibilities, data cleanup, integrations, third-party costs, environments, QA depth, accessibility and security work, training, launch support, warranty or defect handling, ongoing maintenance, and change process.
A fixed price can be useful when the evidence supports a stable scope. A phased or time-based model can be more honest when important dependencies remain unresolved. Neither model removes risk by itself. The commercial structure should expose uncertainty and define how decisions are made when evidence changes.
A hypothetical partner comparison
The following is a hypothetical example, not an AP Works client project, project result, price, or timeline claim.
A business wants a customer and staff portal for service orders. Customers submit requests and files; staff review them; a finance role approves adjustments; the billing platform remains the source of truth; and failed synchronization must enter a visible recovery queue.
Partner A presents a polished portfolio and a feature estimate. Its proposal says it will add “user roles,” “billing integration,” and “full QA,” but does not define record boundaries, workflow states, failure handling, acceptance evidence, or account ownership.
Partner B returns a short state model, flags the difference between customer, reviewer, and finance authority, identifies the billing API as a dependency to validate, proposes a reconciliation queue, lists the acceptance journeys, and shows which repository, cloud, and monitoring access will transfer. Partner B has not designed the whole system for free; it has demonstrated a more traceable way of thinking about the risk.
The scorecard should reward that evidence. It should not automatically choose the candidate with the longest proposal, lowest total, newest stack, or most impressive homepage.
Questions to ask before selecting the company
- Can you restate our representative workflow, including its boundary, authority, exception, and completion evidence?
- Which assumptions must be validated before you can commit to the first release?
- What will discovery produce, who owns those outputs, and what decision ends the phase?
- Who will actually perform product, design, engineering, QA, deployment, and support work?
- How will requirements remain traceable through interface decisions, code, and acceptance tests?
- How will permissions be enforced at the record and action level?
- How will each integration handle limits, duplicates, failures, reconciliation, and provider changes?
- Which security, privacy, and accessibility requirements apply, and what evidence will prove them?
- Which environments, accounts, repositories, licences, and third-party services are included?
- What is the release, rollback, monitoring, backup, incident, and post-launch support plan?
- What will the business receive so another qualified team can operate and extend the application?
- Which assumptions, exclusions, client tasks, recurring costs, and change rules affect the estimate?
Sources and further reading
- NIST SP 800-218: Secure Software Development Framework (SSDF) Version 1.1
- OWASP: Application Security Verification Standard
- W3C: Web Content Accessibility Guidelines (WCAG) 2.2
- Office of the Privacy Commissioner of Canada: Design with privacy in mind
Explore how AP Works scopes and builds custom web applications and business systems, or brief an application project →