July 21, 2026

State Education Agency Resource Guide: How States Can Drive Procurement of Effective AI Tools

Share this article

#2 IN THE SERIES

Welcome to the “Signals of Quality” series — a look at what separates genuinely effective AI-powered ed tech from flashy technology, and how state leaders and advocates can build procurement practices that put student learning first. Given the growing skepticism of artificial intelligence (AI) and ed tech, how should K-12 education leaders think about what “good” tools look like? Safety-oriented frameworks for AI tools abound, but little exists on assessing efficacy and quality for AI tools used by students and teachers in classrooms. Building on early research, this series 1) surfaces five signals of quality in AI-powered tools, and 2) identifies concrete steps state leaders and advocates can take to prioritize efficacy and learning throughout a procurement ecosystem. These insights arise from the AI Policy Hub, a partnership between Bellwether and PIE Network to connect advocates with resources, support, and national education experts.

 

Procurement is a powerful mechanism state leaders can use to ensure that any AI-powered ed tech tools purchased for K-12 classrooms are safe and effective. Yet formal efficacy research is slow, few independent sources evaluate AI tools, and state education agencies (SEAs) typically lack the technical capacity to assess vendor claims in depth. In our second installment of the Signals of Quality series, we focus on actions states can take to support efficacy-oriented procurement of AI tools.

This is not a comprehensive guide; there are other aspects of ed tech procurement — such as safety, interoperability, or cybersecurity — not covered here but still critical to evaluate during procurement. However, the field is early in its thinking about efficacy and quality. The first installment elevated five signals of high-quality AI tools; this piece highlights how states can incorporate those quality signals into procurement processes.

 

Quality Signal What States Should Do
Signal 1: Emphasis on learning outcomes, not technology features.
  • Align on which outcomes matter most before writing a Request for Proposal (RFP). Success metrics — whether they are academic outcomes, durable skills, student behavior or well-being, teacher development, etc. — should anchor and drive every aspect of a procurement process. Defining them up front ensures that what follows aligns with target priorities and outcomes.
  • Require vendors to name the specific instructional problem to be solved. It’s easy to get swept away by generative AI’s capabilities, but high-quality tools solve problems, not rely on flashiness. A vendor should be able to say, “This tool builds student fluency in math by identifying misconceptions using AI,” not just “It uses generative AI to personalize learning.”
  • Demand meaningful efficacy measurement anchored to known benchmarks. Existing frameworks — such as the Every Student Succeeds Act (ESSA)’s Tiers of Evidence framework — give everyone familiar language and common ground on which to work. For rapidly iterating AI tools, the “promising evidence” and “demonstrates a rationale” tiers allow greater flexibility without compromising on evidence. The State Educational Technology Directors Association (SETDA)’s “Evidence-Based” indicator provides model questions aligned to the ESSA Tiers, as well as a list of third-party validators, while Opportunity Labs’ “Evidence” benchmark offers sample procurement language.
  • Consider outcomes-based contracts. These contracts are helpful for defining success metrics up front, establishing data-sharing agreements, and requiring disaggregated analyses for students (e.g., students with disabilities, English learner students, and students performing below grade level). Effective contracts also come with implementation conditions to ensure fidelity.
Signal 2: Productive struggle as a primary pathway for learning.
  • Ask vendors to describe how their tool preserves productive struggle and student agency. Technology shouldn’t always make things easier; productive struggle is critical for long-term learning. Vendors should be able to provide specific examples of guardrails against instant-answer behavior, over-scaffolding, or bypassing effort.
Signal 3: Sound pedagogy and coherence with existing instructional practice.
  • Interrogate how a tool would operate within a classroom. Is the primary user a teacher or a student? Is it used as practice, as part of introducing a concept, or as homework? Is it complementary to a teacher’s expertise or supplementary? This thought exercise forces both states and vendors to consider what exactly a tool adds to a student’s learning experience.
  • Reward alignment with existing high-quality instructional materials (HQIMs). Many states already have standards or guidance on HQIMs, and integrating those creates greater incentives for overall coherence than treating AI tools as standalone products. HQIM requirements can also be paired with interoperability standards to ensure the tool integrates with existing assessment and student information systems.
Signal 4: Technical configurations designed to maximize quality.
  • Invest in capacity building for procurement and review staff. AI tools and infrastructure will continue to evolve, and developing internal knowledge and capacity now will pay off in stronger decisions. In light of hiring and fiscal constraints, states can build expertise through technical advisory committees or by forming partnerships with higher education institutions. 
  • Use interviews to dig deeper into a tool’s design. Vendors constantly have to weigh trade-offs across accuracy, latency (speed), cost, and privacy, but these aren’t always visible in marketing materials, short demos, or RFP responses. Interviews allow SEA staff to ask vendors to explain not just the design choices made, but why, and what other options were considered. 
  • Probe how vendors refine and guard against low-quality AI responses. There are many ways developers can improve the quality of their tools’ outputs: context engineering, fine-tuning for specialized use cases, or retrieval-augmented generation (RAG), which grounds responses in vetted educational content rather than the model’s general training. RAG in particular is a useful signal because it indicates the vendor has invested in curating high-quality reference content, reducing hallucination risk, and ensuring curriculum-aligned instruction.
  • Require disclosure of benchmarking and evaluation results. The presence (or absence) of this infrastructure reveals a vendor’s commitment to ensuring high-quality AI outputs. Standardized vendor templates could surface how the vendor develops rubrics, runs regression tests when models or prompts are updated, or validates AI outputs against human expert reviews.
Signal 5: Attention to market sustainability and long-term planning.
  • Require financial disclosures and contingency planning provisions. States need to know what happens if the vendor exits the market mid-contract. Disclosures could include funding sources (e.g., venture capital, philanthropy, fee-for-service) and a contingency plan that outlines transition support, data portability, and any continued access to the tool’s outputs.
  • Build data ownership and portability clauses into contracts. Data ownership clauses are already standard practice in many contracts, but AI tools heighten their importance. Beyond student data, these tools generate new artifacts — such as student- or teacher-AI interaction logs — that are increasingly valuable for evaluating tools, training future models, and informing instructional decisions. Losing ownership of these artifacts could lock districts into a single vendor or inadvertently endanger student privacy. Contracts should specify what will happen to these data and how student information will be protected if vendors shut down or are acquired.
  • Leverage state-vetted lists and master contracts to reduce districts’ risk. State-negotiated agreements with vetted vendors can streamline district procurement while adding protections that individual districts may struggle to get independently.

 

 

Four High-Leverage Policy Instruments

While districts execute contracts, states shape the broader procurement environment. Across the specific actions above, a few policy levers consistently stand out as effective and efficient ways state leaders can change procurement to better assess AI tools:

  1. Build expertise and technical capacity. Examples include establishing technical advisory committees, engaging AI researchers and data scientists at local universities as third-party reviewers, or facilitating communities of practice or statewide collaboratives. Expertise is a long-term investment that will continue to pay dividends through smarter decision-making.
  2. Shape market demand and vendor behavior. Many ed tech providers are for-profit organizations; as a result, the macroeconomic environment plays a strong role in how tools are designed, built, evaluated, and marketed. States can influence this environment indirectly through standardizing procurement with model disclosure requirements, review criteria, or contract clauses; creating shared statewide accountability standards; or publishing “vetted provider” lists. States can also directly participate in the market through establishing or joining multistate or multidistrict purchasing coalitions, or negotiating master contracts with bulk rates to secure lower per-district pricing and more predictable costs.
  3. Support public evaluation infrastructure. Measurement is hard for AI tool developers, but public support for wider infrastructure initiatives can accelerate evaluation cycles and help policymakers identify effective tools faster. Concretely, states could build public datasets and evidence bases by using pilot programs or encouraging data-sharing agreements among districts, vendors, and researchers.
  4. Continue evaluating tool performance post-procurement. Ongoing evaluation matters more for AI tools than for traditional ed tech. Models update, vendors iterate, and how a tool performs can change materially mid-contract in ways that pre-deployment review could not anticipate. States should consider setting up or encouraging outcomes-linked contracting structures and finding an appropriate cadence for revisiting and elevating reporting results for public transparency.

 

The pace at which AI tools change after deployment, the new categories of data they generate, and the relative immaturity of the vendor market all distinguish these tools from prior ed tech trends, and districts need support to navigate this new landscape. The path is clear for robust state involvement — state procurement and use of AI tools are exempt from federal preemption efforts — but the field’s thinking on AI efficacy frameworks is still nascent. The actions and policy levers outlined here are a first step to help states think about what “quality” means and looks like in AI-powered tools.

More from this topic

Processing...
Thank you! Your subscription has been confirmed. You'll hear from us soon.
ErrorHere