By Abhishek Gang Deb
Key Takeaways
- A useful starting point for applied AI is a specific customer problem and data the application already has permission to use.
- Apple Natural Language can help process transaction descriptions, but tokenization alone does not classify purchases or predict customer behavior.
- A small, explicit category model can connect recent purchases to relevant offers and provide a reason for each recommendation.
- Eligibility, expiry, and confirmed rewards require authoritative business data; they cannot be inferred from text similarity.
- Recommendation quality and business impact need separate evaluation. More visits to an offers screen do not establish additional customer savings or merchant profit.
Moving From AI Tools to a Product Feature
My interest in AI grew as I began using coding assistants. Tools such as GitHub Copilot and Devin made me think differently about my development workflow, but I also wanted to understand how to build an intelligent feature inside an application.
As an Apple developer, I started with familiar ground: Apple’s Natural Language framework. I read about natural language processing, tokenization, word embeddings, and contextual embeddings. The practical question was what I could do with these capabilities in a product I already understood.
I was working on a large banking application. Two existing features caught my attention: recent transactions on the account details screen and a separate collection of merchant offers that customers could activate. The transactions described where customers had recently spent money. The offers represented possible benefits for future eligible purchases. I wanted to connect those two experiences.
That became the basis of a feature I built and demonstrated to the client. In this article, I explain the idea and present a simplified implementation that makes the category matching and ranking decisions easy to inspect.
The Swift examples are a teaching implementation created for this article, not extracts from the bank’s production code. They use synthetic transaction descriptions and a small English-language rule set. The later discussion of evaluation and production safeguards describes recommendations for this design, rather than asserting that every technique was part of the original release.
Connecting Transactions to Relevant Offers
Consider someone whose recent transactions include ride-hailing services and public transit. Transport offers may be more relevant to that customer than a generic selection of promotions. For someone who frequently visits restaurants and cafes, dining offers may be a better starting point.
The idea was to summarize those recent activities into a few categories and match them to categories already assigned to offers. A transaction might contribute to transport or dining; an eligible offer in the same category would become a candidate for recommendation.
This is an estimate of relevance based on recent activity. It does not establish what a customer intends to buy next, and it should not present a category score as a probability of purchase.
My original thinking started with the transaction list visible on the account details screen. Architecturally, the right input is the transaction data that supplies that screen. Reading rendered labels would make the recommendation depend on formatting, pagination, localization, and which rows happen to be visible. A recommendation service should instead receive a defined set of transaction records from the application’s data layer.
For the example below, the flow is straightforward: transaction records become category counts, category counts become a recent activity profile, and that profile ranks eligible offers. The screen displays the results and their explanations. The classification and ranking can be exercised without loading the screen.
Understanding the Role of Natural Language
Natural language processing is a broad field concerned with processing human language. Apple’s Natural Language framework provides several relevant building blocks, each with a different responsibility.
NLTokenizer divides text into units such as words or sentences. It can help turn a merchant description into tokens that an application can inspect. The tokenizer does not supply a banking transaction taxonomy or automatically know that a particular merchant belongs to dining.
NLEmbedding represents strings as vectors and supports similarity comparisons. This can be useful when exploring related words or descriptions, although proximity in an embedding is not proof that two purchases belong in the same business category.
NLContextualEmbedding produces representations that depend on surrounding words. Apple describes its use in tasks such as text classification with Create ML. A merchant classification feature would still need a classification strategy, evaluation data, and appropriate model assets.
For a first implementation, an explicit rule set offers a useful baseline. It allows a developer to see exactly why a transaction received a category. In the example here, Natural Language supplies tokenization; application rules supply the category; frequency supplies the ranking signal. There is no custom trained model or generative AI in this sample.
Where trusted merchant identifiers or existing category metadata are available, I would use those before interpreting a free-text description. Text processing is most useful for records whose structured information is incomplete.
Classifying Transaction Descriptions
Transaction descriptions are short and messy. They may contain merchant names, payment processor prefixes, branch identifiers, or reference numbers. A small difference can matter: a ride-hailing purchase and a food delivery purchase may share the same parent brand.
The first example normalizes case and diacritics, tokenizes the description, and checks a few explicit aliases. Matching complete token sequences avoids matching a short alias inside a longer word. Unknown descriptions return nil.
import Foundation
import NaturalLanguage
enum SpendingCategory: String, Hashable {
case transport, dining
}
func category(for description: String) -> SpendingCategory? {
let text = description.folding(
options: [.caseInsensitive, .diacriticInsensitive],
locale: Locale(identifier: “en_US_POSIX”)
).lowercased()
let tokenizer = NLTokenizer(unit: .word)
tokenizer.setLanguage(.english)
tokenizer.string = text
let words = tokenizer.tokens(
for: text.startIndex..<text.endIndex
).map { String(text[$0]) }
let normalized = ” ” + words.joined(separator: ” “) + ” ”
let rules: [(String, SpendingCategory)] = [
(“uber eats”, .dining),
(“uber trip”, .transport),
(“lyft”, .transport),
(“bart”, .transport),
(“dominos”, .dining),
(“coffee”, .dining)
]
return rules.first {
normalized.contains(” ” + $0.0 + ” “)
}?.1
}
These aliases illustrate the mechanism, not a production merchant directory. Real descriptions need tests for conflicting names, processor prefixes, punctuation, abbreviations, and regional variations. Even a complete token match can be wrong: the word “coffee” in a merchant name does not guarantee a cafe purchase.
An unknown category is a useful result. Forcing every description into one of two buckets would make the profile look complete while hiding classification errors. The application should preserve unknown records for coverage measurement and avoid generating specific explanations from them.
The function creates a tokenizer for each invocation. Apple’s threading guidance for NLTokenizer requires serialized access to a shared instance, or separate instances for independent execution contexts. For a larger batch, a worker can reuse its own tokenizer sequentially without exposing that mutable object to concurrent callers.
Building a Recent Activity Profile
The phrase “spends more” needs a precise definition. It might mean the highest total amount, the most purchases, or the most recent activity. These choices can produce different recommendations. A single large purchase can dominate an amount-based profile, while frequent smaller purchases may better reflect a repeated habit.
This example uses purchase frequency within the preceding 30 days. That is a transparent starting rule, not a claim that 30 days is optimal. The service supplying the records must identify posted purchases and reconcile refunds, reversals, and pending-to-posted transitions according to the bank’s transaction semantics.
struct Transaction {
let id: String
let description: String
let postedAt: Date
let isPostedPurchase: Bool
}
func purchaseCounts(
transactions: [Transaction],
now: Date
) -> [SpendingCategory: Int] {
let cutoff = now.addingTimeInterval(-30 * 24 * 60 * 60)
var seen = Set<String>()
var counts: [SpendingCategory: Int] = [:]
for transaction in transactions {
guard transaction.isPostedPurchase,
transaction.postedAt >= cutoff,
transaction.postedAt <= now,
seen.insert(transaction.id).inserted,
let value = category(for: transaction.description)
else { continue }
counts[value, default: 0] += 1
}
return counts
}
The identifiers must refer to canonical transactions within the supplied account scope. The set prevents duplicate records from increasing the counts, but it cannot reconcile two different identifiers for the same underlying purchase. That work belongs upstream.
Suppose a synthetic input contains three classified transport purchases and two dining purchases. Transport contributes 60 percent of the classified activity and dining contributes 40 percent. An unrecognized sixth purchase contributes to neither category. Those percentages describe the classified subset, so the fraction of unknown records matters when judging whether the profile is representative.
For accounts with very little usable history, a generic list of eligible offers is a reasonable fallback. A production design should also consider limiting repeated purchases from the same merchant so one narrow pattern does not dominate the entire recommendation list.
Ranking Offers After Checking Eligibility
An offer category makes matching possible, but it does not establish eligibility. Card restrictions, activation requirements, merchant exclusions, campaign dates, and other terms still apply. The offer service should resolve these rules and supply authoritative eligibility and validity information.
The next example ranks only eligible offers that are currently valid and whose categories occur in the profile. It uses the category’s share of classified purchases as a relevance score. For equal scores, earlier expiry comes first, followed by an identifier to keep ordering deterministic.
struct Offer {
let id: String
let title: String
let category: SpendingCategory
let validFrom: Date
let expiresAt: Date
let isEligible: Bool
}
struct RankedOffer {
let offer: Offer
let score: Double
let explanation: String
}
func rankOffers(
_ offers: [Offer],
counts: [SpendingCategory: Int],
now: Date
) -> [RankedOffer] {
let total = counts.values.reduce(0, +)
guard total >= 3 else { return [] }
return offers.compactMap { offer -> RankedOffer? in
let count = counts[offer.category, default: 0]
guard offer.isEligible, offer.validFrom <= now,
now < offer.expiresAt, count > 0
else { return nil }
return RankedOffer(
offer: offer,
score: Double(count) / Double(total),
explanation: “Based on \(count) recent ”
+ “\(offer.category.rawValue) purchases.”
)
}.sorted {
if $0.score != $1.score { return $0.score > $1.score }
if $0.offer.expiresAt != $1.offer.expiresAt {
return $0.offer.expiresAt < $1.offer.expiresAt
}
return $0.offer.id < $1.offer.id
}
}
The minimum of three classified purchases is an illustrative product threshold. It is not a confidence guarantee and should be tested with representative data. The sample also assumes unique offer identifiers and defines expiry as an exclusive instant. A backend must translate campaign dates and time zones into that contract consistently.
Expiry breaks ties; it does not turn an unrelated offer into a relevant one. Likewise, category matching does not prove an offer will save the customer money. The interface still needs to display the terms, and eligibility must be rechecked when the customer activates or uses the offer.
Bringing the Recommendation Into the Account Experience
My interaction idea began on the account details screen. An animation around recent transactions would draw attention to the connection with available offers. A bell-style entry point would then take the customer to the offers section.
The explanation was central to that experience: “Based on your recent transactions, these offers may be relevant to you.” A more specific explanation can mention recent transport or dining purchases when the underlying classifications support it.
The destination also needs to distinguish different kinds of information. Recommended offers describe possible future relevance. Previously earned benefits describe confirmed outcomes. Expiring offers describe a time constraint. Combining them into one undifferentiated message would make it difficult for customers to know what had already happened and what still required action.
In particular, a similar past transaction does not prove that an offer benefited the customer. A claim about savings already received should come from confirmed reward or redemption records. Otherwise, the interface should describe potential relevance without implying a completed benefit.
For a production interface, I would keep the entry point dismissible, provide an accessible label, and respect Reduce Motion. A persistent animation or urgent notification treatment could distract from the account information the customer came to see. The feature should remain useful when animation is disabled.
Keeping the Feature Predictable in Production
The sample classification and ranking code can execute locally without sending descriptions to an external inference service. That is a property of this implementation, not a claim that the whole banking application operates offline or that the original production feature used this exact architecture.
A production service should process a bounded batch away from the UI’s critical rendering path and publish only the result needed by the view. Cache invalidation should follow changes in transactions and offers. Switching accounts or signing out must also invalidate derived profiles so recommendations cannot cross account or session boundaries.
Existing access to transaction data does not by itself settle whether every personalization use is appropriate. The implementation should follow the product’s established permission and retention rules. Raw descriptions should not appear in analytics events merely to explain ranking; category-level diagnostics may be enough.
Finally, recommendation failure should not prevent account details from loading. Empty history, missing categories, stale eligibility, and unavailable offers are normal states. An ordinary offers list or no recommendation at all is preferable to a misleading personalized message.
Evaluating Relevance and Business Results
When I demonstrated the feature, the client responded positively. It was subsequently used in the live application and, in my experience, encouraged customers to explore the offers section. This account describes that outcome qualitatively; it does not include a measured uplift or a quantified merchant profit claim.
To evaluate a system like this more rigorously, I would begin with classification quality. A reviewed dataset should include unfamiliar merchants and ambiguous descriptions as well as easy matches. Per-category precision, coverage, and the unknown rate reveal different problems: a system can be precise on the few records it recognizes while failing to classify most purchases.
Next, I would compare the recommendations with a simple baseline, such as the existing eligible-offer ordering. Offer views, activations, and confirmed redemptions should remain separate events. Each measurement needs a defined denominator and observation period. A controlled experiment, where feasible, helps distinguish the feature’s effect from changes in campaigns or customer activity.
Performance also needs device measurements: processing time for a realistic batch, memory use, and any effect on account-screen responsiveness. None of those results should be inferred from the small size of the example code.
Merchant profit is a further business question. Additional clicks, activations, or even purchases do not alone establish incremental profit. That claim would require appropriate commercial data and a comparison that accounts for what would have happened without the feature.
Where I Would Explore Machine Learning Next
The rule set provides a baseline against which to evaluate more advanced approaches. If its coverage is too limited, embeddings or a supervised classifier may help with descriptions that do not match known aliases. The decision to add a model should follow measured shortcomings and representative examples.
I would evaluate embedding similarity against labeled merchant descriptions before using it to assign categories. Generic language similarity can miss distinctions that matter to the business, especially in short strings with little context. Contextual embeddings also require suitable model assets and an explicit strategy for converting their output into a category.
A learned classifier should be able to abstain when its evidence is weak, and it should be compared with the same baseline on held-out data. Category prediction would still remain separate from offer eligibility, ranking policy, and confirmed rewards.
For me, this project connected learning about AI with a familiar engineering responsibility: deciding what a feature should do with the data available to it. The useful part was making a connection customers could understand, then defining exactly what that connection did and did not establish. That is where I would encourage another iOS developer to begin.
References
- Apple Natural Language framework
- Apple NLTokenizer documentation
- Apple NLEmbedding documentation
- Apple NLContextualEmbedding documentation
Abhishek Gang Deb is a technical architect focused on iOS engineering, production reliability, Swift concurrency, and large-scale mobile application architecture. His work includes debugging complex production incidents, improving architecture boundaries, and building safer abstractions for mobile engineering teams.
