Saturday, June 21, 2008

Managing Organizational Risk II: Controlling The Systems Flow Process

A significant source of organizational risk is not having sufficient control of the activities and processes within your department or organization. In today's fast paced and rapidly changing environment, significant parts of any process have to be broken down and delegated to various people, often with disparate sets of skills and experience levels. Having several people simultaneously making changes to specific parts of a process in order to make the entire process responsive to a sudden change in the environment can be quite a difficult task; and if not handled properly can (1) result in duplication and waste of valuable work hours, (2) create confusion and breed hostility in the workplace, (3) cause the organization to fail to respond to the crisis that prompted the process change exercise, or even (4) muddle or damage the existing process. The steps below provide a generic approach to manage changes in the organizational systems flow process:

Figure 4.1 : Controlling the systems flow process.


Step 1
The first step is to baseline the processes within your organization. Accumulate documentation on existing processes. If documentation is lacking, interview the relevant people and ask them to document their processes. Make sure that these documents are complete and detailed enough for other people to execute. Create a master document of how these processes work together(i.e., The Master Strategy document).

Step 2
Any change to the process (or part of a process) should only be initiated via a written request. Subsequently, The request's feasibility and viability should be evaluated after it has been received, granted that requests for certain processes would go through a more rigorous evaluation than others. It is not a good idea to make changes to a process, or part of a process, without some sort of a document trail.

Step 3

After the request has been received and evaluated, the next step is to determine whether your department has a pre-existing methodology to handle the change request. Often many of the change requests received are of a recurring nature, your organization may have an existing methodology to handle the change. This is another reason why documentation is important.

Step 4a
If a methodology exists for the change request, you need to verify whether it is sufficient for the request. You may need to make minor modifications to the methodology to adhere to a specific request.

Step 4b
If a methodology does not exist, you need to develop a new one to manage the change request. When developing the new methodology, you need to be cognizant of the fact that it may affect methodologies belonging to other systems, and even require you to develop additional methodologies. Also, you will need to design tests for the newly developed methodology to ascertain that it provides accurate and consistent results.

Step 5
Once you have tested and are fairly confident about performance of the new methodology, the next step is to execute it within the process and validate the outcome. Inform other stakeholders about the changes that have been implemented within the process.

Step 6
After the change request has been completed, the next step is to update the Master Change document. The Master Change document maintains a continuous list of changes that are made to any of the processes for future reference.

Step 7

Determine whether the change request should be part of a broader strategy.If so, update your Master Strategy document.

Step 8
The final step is to develop and execute a post implementation validation plan for the new strategy. The purpose of this plan is to make sure that the outcome of the new changes remain accurate and consistent over time.

Sunday, May 11, 2008

Tools of the Trade III: Simple Data Validation Procedures

The reliability of your analytical model depends to a great extent on the quality of its underlying data. Unfortunately, bad data is a prevalent problem for most corporations. A survey by Gartner found that at least 25 percent of all data for Fortune 1000 firms were bad or corrupted. Based on my experience, carrying out several simple procedures after reading in the data will let you track and rectify many of these data quality problems, even if you don't get 100 percent accuracy. If you are getting your data from an outside source, it is a good idea to share your findings from the data validation procedures with it. The people providing you the data may be able to tell you whether your preliminary findings make sense or needs further investigation. Carry out these procedures, even if you have assurances that the data is good and complete. These simple - yet effective - procedures include the following:

1. Prior to performing any analyses, pull out the first and last 15 observations of the data and look for any obvious inconsistencies. Examine columns with blanks and null values. Check and see whether the values for a variable make sense (i.e., a column that should have numeric values, such as price, does not contain numeric values).

2. If using SAS, perform a contents procedure on the dataset to see if all the variables are accounted for in the proper format. Date and Time are often the trickiest variable. Make sure whether they are in time, timestamp or character format. IDs are another problematic variable. Generally, they should be in character format (even if they are numbers), but often exist in numeric format.

3. Verify the number of records you read into your program. Investigate if you have more (often resulting from trailing blanks) or less records (records getting dropped by the read in program) than anticipated. Trailing blanks are not harmless and may cause your program to malfunction during later stages of your analyses.

Try not to eliminate any records. Save the bad records in separate datasets for further analyses. These records often provide valuable clues for future QA strategies.

4. Make sure the data falls within valid ranges. For instance, if you know the age of people in the dataset should not be less than 18 years or more than 65 years, all the DOB should be within that range. If you know that the highest and lowest unit prices in your dataset are $100 and $0.10 respectively, then seperate and study records where unit prices are above $100 or less than 10 cents.

5. Perform a frequency on key variables. A frequency on dates could show records missing for an entire day. Also, pay attention to dates equavalent to default dates (i.e., Jan 1, 1900 for Excel and Jan 1, 1960 for SAS). Make sure there are no blanks when there should be null values, and vice versa.

6. Randomly pull 15 records from the entire dataset and ascertain that there are no egregious errors.

7. Read out the records with the 10 highest and lowest values for key variables.

8. Output the mean, median and mode for key variables.

Save your data validation results. Comparing the different findings over time should enable you to identify and correct data problems.

Using SAS to compare two datasets

You can use the SAS compare procedure to validate a dataset by comparing it with a baseline dataset. Use the following code to compare two identical datasets:

proc compare base=indat1 compare=indat2 OUTSTATS=outdat1 noprint;
run;

This is helpful when you need to modify an existing SAS program, but do not want the modifications to change the output.

Tuesday, April 15, 2008

Managing Organizational Risk I: How To Deal With Transactional Fraud

If you read the newspapers or watch the news, you are probably aware that transactional fraud has been increasing exponentially and will continue to do so in the near future. According to a Nelson report, financial institutions and online merchants have lost over $ 1.2 billion in 2005 alone. Transactional fraud, as opposed to regular fraud, occurs when a credit card holder or online account holder denies authorizing or engaging in transaction(s) involving his/her credit card or account. This narrow definition excludes other instances of fraud from the purview of “transactional fraud”. If a customer, for instance, authorizes a transaction to purchase an item from a vendor, and the vendor intentionally delivers a lower quality product to the card member, that particular action will not fall within our definition of transactional fraud. Such cases should be referred to the organization's disputes division, or the affected person should pursue legal action against the vendor. In order to be classified as transactional fraud, the customer must deny authorizing or engaging in that particular transaction with the vendor.

Scenario #1: Customer calls to inform that he entered into a contract with a vendor to purchase merchandise. Vendor charges bill to customer’s credit card but does not deliver merchandise. Since customer authorized the transaction, this case does not fall within the purview of transactional fraud. Customer is referred to the billing disputes department.

Based on this narrow definition of transactional fraud - it becomes critical for companies to properly establish the identity of the customer - often over the phone. As a result, any strategy to deal with transactional fraud often includes a investigator or telephone rep who calls up the customer and tries to determine whether that person is the actual card holder or online customer. Helpfully, there are various identity verification tools in the market, and an investigator would avail several of these tools in conjunction to design tests that the customer needs to pass in order to prove his or her identity. Credit rating bureaus, such as Experian and Transunion, offer several of these verification tools, which include:

  • Fast Data: Allows investigators to verify past and present addresses and phone numbers for a customer.

  • First Pursuit: Provides data from the credit rating bureaus for a credit card holder.

  • GUS: Allows investigators to view application and credit bureau data for a credit card holder.

  • MARS: Allows investigators to view information on other accounts linked to a credit card holder.

  • VERID: Enables an investigator to establish the identity of a credit card holder by asking a series of questions that only the card holder would know.

  • VRU: Enables an investigator to see all the phones numbers that were used to inquire about a particular account or transaction.

In addition, many internet security firms allow companies to trace computer IP addresses that may be used commit transactional fraud via the internet. Your company should also have the ability to allow investigators query information from your company databases as well as create and maintain memos for accounts that have been "touched". Often, past information pertaining to a specific account becomes crucial to identify fraud or foul play. There are also software applications in the market, such as Falcon offered by Fair Isaacs and Fraud Management offered by SAS, that can analyze millions of transactions in real time and generate alerts on suspicious account activity based on predetermined rules.

Depending on how the alerts are generated or brought to your attention, the procedures for managing transactional fraud can be grouped into two main strategies: (1) Inbound and (2) Outbound. Both these strategies are further discussed below:

Inbound Fraud Strategies

These strategies are employed when a customer calls your organization to complain about potential fraud activity on his account. Companies often set up several groups to manage the flow of inbound fraud cases:

The Primary Inbound Fraud Unit: The primary inbound fraud unit is the gateway for handling potential inbound fraud cases. These cases may arrive at the inbound unit from different departments of your organization, such as Customer Services, Billing Disputes, Collections, etc. A customer may personally call your fraud department to report fraudulent activity on her account. Many outbound fraud strategies may place a temporary block on the account that may prompt the customer to call your company. After receiving the call, an investigator performs a quick series of verification tests on the customer to ascertain her identity utilizing one or more of the validation tools described above. If the investigator ascertains that no transactional fraud was committed, he will fill out a memo explaining his findings and close the case. Future investigators who pull up the account would access to the contents of this memo. If the investigator, however, suspects that transactional fraud was committed or the customer fails the verification tests, the account is temporarily blocked and the case is referred to the secondary inbound fraud unit for further investigation.

Scenario #2: Customer calls bank because she can’t use her credit card. Investigation reveals a temporary block placed on the card because a large jewelry purchase in the Bahamas was charged on it. Investigator verifies card holder’s identity and account information, and removes temporary block. Investigator writes up her findings in a memo that will be available to future investigators who pull up the account.

The Secondary Inbound Fraud Unit: Potential fraud cases that have not been resolved at the primary inbound unit and require further investigation are referred to the secondary inbound fraud unit. The secondary inbound fraud investigator, upon receiving the case, would read the memo written by the primary inbound investigator, and perform further verification tests on the customer. Secondary inbound investigators are more experienced people who have doing this kind of work for a while. The investigator would then place a call to the customer to gather more information on the incident. If the investigator resolves the case, he removes the block from the account and records his findings in a memo. If the investigator suspects the case involves fraud, he places a permanent block on the account and fills out a fraud application. The case then gets referred to the investigations division. Permanent blocks are more difficult to remove than a temporary block.

Scenario #3: Customer calls bank to inquire about a credit card statement she received for a card that she did not apply for. Investigations reveal suspect(s) intercepted credit card offer letter, opened account and transferred large amounts of cash to their account. Customer is advised to contact credit rating agency and file a police report. Investigator puts a permanent block on the account and fills out a fraud application. Customer has no liability for money stolen.

Scenario #4: Customer was subscriber to a large Internet Service Provider (ISP). Customer claims to have called ISP to cancel service two years ago. ISP did not cancel service and was billing customer for two years. The charges appeared on the customer’s credit card statement for the past two years. Customer called ISP to get charges reimbursed. ISP refused to reimburse charges. Customer called bank to claim credit card fraud. Case is not fraud because customer had initially authorized charges and was not diligent in ensuring that the ISP had actually canceled service. The case is referred to the billing disputes department.


Figure 3.1: How a large commerical bank handles inbound fraud.

Outbound Fraud Strategies

These strategies are employed to deal with transaction fraud when the potential fraudulant activity is brought to your attention other than the customer calling your organization. Potential fraud activity can be brought to your attention from transactions pertaining to "compromised" accounts (i.e., accounts that are known to be stolen), credit rating bureaus, electronic fraud detection systems such as Falcon, etc. In many cases, accounts that are potential fraud cases have a temporary block placed on them. As with the inbound fraud strategies, the outbound fraud strategies can consist of various stages and groups to better classify and resolve the potential fraud cases. For example, you can have a primary outbound fraud unit that analyzes the preliminary outbound fraud cases and routes them to a group that specializes in a specific fraud type, such as Takeover or NRI.

The Primary Outbound Fraud Unit: After the investigator receives the potential fraud case, he performs a series of verification tests on the customer and account using the available validation tools. One of these tests often involve a process called Trending, where the investigator tries to determine at least three matches for a customer from the customer's personal information (i.e., phone numbers, SSN, driver’s license, etc.) stored in different systems. The investigator would then place a call to the customer. The investigator may request additional documentation from the customer to validate his identity. If the investigator is able to determine no transactional fraud was involved, he would remove the temporary block, record his findings in a memo, and close the case. If the case involves transactional fraud, the investigator will check the outstanding balance on the account. If the outstanding balance is zero, the investigator would put a permanent block on the account. If the outstanding balance on the credit card is positive, the case is referred to the secondary outbound group for further analyses. If the investigator is unable to get hold of the customer, he leaves a message and a contact phone number. When the customer calls, her case is handled by the inbound group, which has access to the outbound investigator's memos pertaining to the account.

Because you do not have an impatient customer on the line and more time to research the potential cases, outbound fraud strategies can also be tailored to respond to specific fraud types, such as:

i. Fraud Applications: Cases where perpetrators steals a victim’s personal information and use it to open an account.

Scenario #5: Credit Card issued from bank in high risk area (this particular branch has issued credit cards to fraudulent applicants in the past) and customer fails VERID test. Since account has positive balance, case is referred to the secondary outbound unit for further investigation.

ii. Takeover: Cases where the the suspect steals the identity and account belonging to a customer and commits fraud. For credit cards, the phone number used by a suspect to activate the card or inquire about the account is used to determine whether fraud is being committed.

Scenario #6: Customer lives in Los Angeles, but the call to activate the credit card was placed from Texas. Investigator calls customer, and customer states that he has a cell phone with a Texas area code. Customer passes VERID. Case is not fraud.

Scenario #7: Customer’s ex-husband steals credit card and purchases items from a major retailer. This is takeover fraud because the suspect stole both the customer’s identity and account. Investigator fills out fraud card, puts a permanent block on the account, and updates the card member’s information in the system.

iii. NRI (Never Received Issue): Cases where the perpetrators intercept the credit card before it reaches the customer and commits fraud.

Scenario #8: Card member applied but did not receive credit card. Since credit card account has zero balance, the investigator puts a permanent block on the credit card account and fills out a fraud card. Investigator also initiates the process to get customer a new credit card account.

Scenario #9: Customer applied for a credit card, but claims to have never received the card. However, the card was activated from his home phone. Customer claims to have not been in his apartment during that period. Investigator requests customer to provide additional documentation to establish identity and presence, and refers case to the Investigations Division.

iv. CNP (Card Not Present): The suspects use credit or debit account information without the physical card being involved, usually through e-mail or other electronic means.

Scenario #10: Customer lists wrong home number, but uses the correct PIN to withdraw more cash than her credit limit. Outbound unit investigator performs verification tests and calls card member. Card member passes VERID. Case is not fraud.

Scenario #11: Electronic surveillence system detects a newly opened account making a $10 purchase at the website of a well known electronics merchant followed by a $6000 purchase within 30 minutes. Investigator unable to get hold of customer using the information provided. Account is blocked. Customer never calls back. Case is considered fraud.

The Secondary Outbound Fraud Unit: Potential outbound fraud cases that have not been resolved at the primary outbound unit and have positive balances are referred to secondary outbound fraud unit. The secondary outbound investigator, upon receiving the case, would read the memos written by the investigators from previous unit(s), and perform further verification tests on the customer and account. If the investigator is able to determine that the case does not involve transactional fraud, he removes the temporary block on the account, makes the necessary credit adjustments, and updates the customer's information in the systems belonging to various credit rating agencies. If the investigator is unable to resolve the case or suspects that fraud is involved, he puts a permanent block on the account, requests the customer for further documentation to establish his identity and sends the case to the investigations division.

Evident from the process described above is that investigators - working both the inbound and outbound strategies - go great lengths to establish the identity of the customer or card holder. By establishing the identity of the customer and ensuring he or she approved the transaction, the company can often defer the monetary loss (based on contractual agreements) accrued on that account or credit card to other entities. Hence, for your company, customers who protect themselves against identity theft are still your best line of defense against transactional fraud.
By taking the following basic precautions, customers can protect themselves against identity theft, as well as greatly enhance the effectiveness of your transactional fraud prevention strategies:
  • Never give out personal information over the phone, over the internet or through the mail unless the customer initiates the transaction or knows who he or she is dealing with.
  • Protect mail. Get incoming mail in a locked mailbox or slot. Take outgoing mail to a postal mailbox or the post office. If mail suddenly stops, go to the post office. Thieves sometimes submit change of address forms to divert mail to their addresses.
  • Check bank and credit card statements carefully. If there are any problems, report these problems immediately.
  • Use a shredder to destroy papers containing sensitive information, such as account numbers, birth dates, SSNs. Destroy all solicitation letters and balance transfer checks sent from credit card companies.
  • Monitor credit reports. There's a website called FreeCreditReport.com that enables people to do that. Or they could contact the three credit bureaus: Equifax (800-525-6285), Experian (888-397-3742) and Transunion (800-680-7289).
The Federal Trade Commission (FTC) website, http://www.ftc.gov/, offers more information about preventing identity theft and what to do if someone thinks his or her identity has been stolen. The FTC's toll free ID Theft Hotline is (877) ID-THEFT (877-438-4338).

Saturday, April 12, 2008

Laying the Analytics Foundation II: Designing a Questionnaire

Designing a successful questionnaire often involves balancing two dueling objectives: (1) getting all the information you need, and (2) persuading your survey participant to provide the information. Ideally, your questionnaire should not have too many questions or be too difficult to fill out; least your participant becomes frustrated and quits, or worse, provides bogus information. At the same time, if you do not get the information you need, the entire exercise becomes a waste of time and resources. Prior to designing the questionnaire and carrying out the survey, it is assumed that you have tried to obtain the information from secondary sources and failed. Gathering information by primary sources, such as a survey, is almost always more expensive than obtaining it through secondary sources.

Before creating the questionnaire, determine exactly what information you need for your analyses. Use short and simple questions to query the information. Avoid using difficult or ambiguous language. The rule of thumb (what I've been told) is that an 8th grade student should be able to read and completely understand the questions. Make a good faith effort to limit the number of questions to what you absolutely need. Provide enough space for respondents to be able to write out the answers.

The next step is arranging the questions according to some logical sequence, to not confuse the participant. If you look at the example below (Figure 2.1), the questionnaire was designed to obtain information on USPS packages transported by rail vans. We decided gathering information on the packages was not enough for our analyses; we also needed information on the rail vans and plants. Accordingly, we divided the questionnaire into three parts. The first part pertained to the rail plant, because that's what the data collector would first encounter. Once the data collector had entered the plant and filled out the necessary data, the next step was finding a rail van. Hence, the second part of the questionnaire involved gathering data on the rail van. Finally, after locating the rail van, the data collector would be able to find the mail packages and fill out the third and final part of the questionnaire.

Carefully decide the type of question to include in your questionnaire. Your questionnaire can consist of Structured questions or Unstructured questions, or both. For structured questions, you essentially know the answers of the questions and force the participant to provide a specific answer. It may be multiple choice, binary (i.e., Yes/No), or inquire for a specific type of information (i.e., Mail Code). The common predicament with structured questions is that you have to know the potential answers in advance. Rarely, you get unknown information with a purely structured questionnaire. Unstructured questions, on the other hand, provides the participant with a free form to volunteer information. Unstructured questionnaires allow you to uncover new details about your test subject. However, you might not get the information you need for your analyses. Although structured questions are great for analytics purposes, use some unstructured questions to provide some flexibility in your questionnaire (See Question #24 in the sample questionnaire).

In our sample questionnaire, you'll see that I tried to fit everything on two sides of a single page. The front page contained the questions, while the back page contained the instructions for answering the questions. This was intentionally done to simplify the job of printing and distributing the questionnaires. I provided a brief purpose so that the data collectors had a broad overview of why we were gathering the information. Each question number in the front page had a corresponding number in the back page that provided the instructions on how to collect the data. Exceptions were highlighted in bold or underlined to draw attention. I also provided hints on where the data collector could find the necessary information. If a data collector - after reading the instructions - had any questions about the survey or procedures, I provided the name and phone number of a contact person to help him/her out. For contingencies where data collectors had to record a lot more data than expected, I provided supplemental questionnaires.

On the bottom right corner of the front page, you'll see a space for processing code. The processing code is used to tag the completed questionnaire after you receive it. It is good practice to save the original paper copies. During later stages of data processing and analyses - if you ever stumble on data that makes no sense - you can use the processing code or tag number to pull up the original questionnaire and see how it was filled out.

Finally, test your questionnaire once it is completed. Give copies to people you know and ask to fill them out. This will allow you to identify and fix any wrinkles you may have overlooked.


Figure 2.1: Both sides of a sample questionnaire

Saturday, December 29, 2007

Tools of the Trade II: Using SAS to Extract the Data

In this post, we will be using SAS to read in the input data from different file formats.

Reading data from a text file
Reading data from a text file is the most basic and popular way to read in data for a SAS program.Usually this is how analysts first learn how to read in data into their programs. It gives you the greatest control while reading in the data, and you can read in as much data as you want. Let's assume that there is a text file named 'hsbctxns.txt' saved in a folder called hsbcdata in your computer's C\: drive. The code for reading that file is given below:

data raw; infile
"C:\hsbcdata\hsbctxns.txt" missover pad lrecl=1000;

input
@1 cardname $8.

@9 merchant $20.

@29 amount 8.
;
run;


If you want greater control over the text file you want to read in, you can utilize the following code. The double questions marks (??) allows the code to read in data even if the format type is different what is specified in your code. In the example below, your code would read in amount values, even if some of them are in character format, although the amount field has been specified as numeric.

data raw;
infile "C:\hsbcdata\hsbctxns.txt" missover pad lrecl=1000;

input
@1 cardnom1 ?? $8.
@8 cardnom2 ?? $8. @;
cardnom2 = cardnom1;
input
@19 merchant ?? $20.

@39 amount ?? 8.
;
run;


To save type typing code and directly import a spaced tab *.txt file to your program, use the following code:

PROC IMPORT OUT= WORK.INDAT
DATAFILE= "C:\REJECT_REPORT.txt"

DBMS=TAB REPLACE;

GETNAMES=YES;

DATAROW=2;

RUN;


To import a delimited *.txt file, use the following code:

PROC IMPORT OUT= WORK.Profit_Rpt
DATAFILE= "C:\Profit_Rpt.txt"

DBMS=DLM REPLACE;

DELIMITER='00'x;

GETNAMES=YES;

DATAROW=2;

RUN;


Reading data from a *.csv file

Now we'll read in data from delimited or *.csv file. The delimiter let's the program know when the current field ends and the next field begins.
This method is advantageous over the previous method because you don't have to specify the position and format for each variable you read. Also, this method lets you read in as much data as you want. Let's assume there's another file named 'deptamt.csv' in the hsbcdata folder in your C:\ drive. Here is how you would read in that data:

data rawcsv;
infile "C:\hsbcdata\deptamt.csv" dlm=',' dsd missover pad firstobs=2 lastobs=1000;

length name $8. type $4. amount 8.2;

input name $ type $ amount;

run;

One problem you might encounter using comma (',') as a delimiter is that numbers (i.e., 23,000) or certain fields (i.e., LastName, FirstName) may have commas within their values that may corrupt your dataset. A less frequently used special character, such as tilda ('~'), as a delimiter can help you avoid this problem.

To import a *.csv file, use the following code:

PROC IMPORT OUT= WORK.Employee_Records
DATAFILE= "C:\Employee_Records.csv"

DBMS=CSV REPLACE;

GETNAMES=YES;

DATAROW=2;

RUN;


Reading data from an Excel file
Now on to reading in data from a Microsoft excel file. This saves the effort of converting your Excel file into another file format and reading it in. However, Excel limits the row and columns of data you can read in. This code lets you read in the data from an Excel file named hsbcfigs.xls:

filename toSAS dde"Excel[hsbcfigs.xls]hsbcstat!r1c1:r100c100" notab;
data rawxls;

infile toSAS dlm='09'x dsd missover;

length name $8 type $4 amount 8.2;

input name$ type$ amount;
run;


To import an Excel file, use the following code:

PROC IMPORT OUT= WORK.indat
DATAFILE= "C:\valid_tests.xls"

DBMS=EXCEL2000 REPLACE;

GETNAMES=YES;

RUN;


Reading data from a DB2 Database
Utilizing the SQL pass thru code below, you can read data directly from DB2 database to your SAS program. The same pass thru code can also be used to read in data from an Oracle database.

proc sql;
connect to db2 (database=app user=appguest password=app123 dsn=core schema=appcore
);
create table testdata as

select * from connection to db2
(select
m.name, m.type, m.amount
from appcore.sales m
where m.salesdate='10JUN2007' and m.amount<=100.00);
disconnect from db2;

quit;


Often, the data for your analytical model would come to you from different sources in various forms. You can use SAS to read in the data based on how the data is made available to you.